MiniMax-M3 / MiniMax-M2.7
MiniMax · 2026-06-01 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
MiniMax-M3
MiniMax M3 被发布文定位为将前沿编码、1M 上下文与原生多模态合一的单模型(Frontier Coding, 1M Context, Native Multimodality - All in One Model),采用 MSA 稀疏注意力并支持思考开关。评测覆盖编码、协作代理、GUI 与多模态:SWE-Bench Verified 80.5、BrowseComp 83.52、OSWorld-Verified 70.06、IMO 2025 35/42。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Coding and Agentic Capabilities · row: SWE-Bench Pro · quote_snippet: SWE-Bench Pro: 59.0%
{
"harness": "Claude Code (scaffolding)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-pro not yet in data/benchmarks.json. Prose bullet in 'Frontier Coding and Agentic Capabilities'; identical 59.0 appears in the results table image and hero chart (two independent vision reads agree). Protocol: internal infrastructure, Claude Code scaffolding, testing logic aligned with official evaluation.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Coding and Agentic Capabilities · row: Terminal-Bench 2.1 · quote_snippet: Terminal-Bench 2.1: 66.0%
{
"harness": "Terminus 2 (scaffolding)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "max output tokens 128K",
"turn_limit": null,
"time_limit": "2 hours",
"run_count": null,
"aggregation": null,
"judge": null
}Prose bullet; identical 66.0 in table image and hero chart. Methodology: sandbox 8C16G, 2h timeout, max output 128K, Terminus 2 scaffolding. Competitor cells GPT-5.5 / Gemini 3.1 Pro / Claude Opus 4.7 are taken from the official Terminal-bench 2.1 leaderboard, others same-infrastructure API runs.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Coding and Agentic Capabilities · row: SWE-fficiency · quote_snippet: SWE-fficiency: 34.8%
{
"harness": "Claude Code (scaffolding)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "2 hours",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swe-fficiency not yet in data/benchmarks.json. Prose bullet; identical 34.8 in table image. Methodology: open-source SWE-fficiency dataset and workflow, sandbox 1C2G, 2h timeout, Claude Code scaffolding, internal testing.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Coding and Agentic Capabilities · row: KernelBench Hard · quote_snippet: KernelBench Hard: 28.8%
{
"harness": "Claude Code",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "benchmark score = mean TFLOPs-ratio over 9 questions",
"judge": null
}Prose bullet; identical 28.8 in table image. Maps to existing benchmark korb (KernelBench), variant Hard. Methodology: Claude Code on NVIDIA Blackwell (CUDA capability sm_120); per-question score = submitted operator TFLOPs relative to hardware theoretical peak; benchmark score = average of 9 questions.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Coding and Agentic Capabilities · row: MCP Atlas · quote_snippet: MCP Atlas: 74.2%
{
"harness": "official MCP Atlas codebase",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "public set",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Gemini 2.5 Pro (aligned with official model)"
}new-benchmark: mcp-atlas not yet in data/benchmarks.json. Prose bullet; identical 74.2 in table image and hero chart. Methodology: official MCP Atlas codebase, Public Set scores with Gemini 2.5 Pro as scoring model, aligned with official model.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Real-World Tasks / Letting M3 Train Models · row: PostTrainBench · quote_snippet: M3 ultimately scored 0.37, slightly below Opus 4.7 (0.42) and GPT-5.5 (0.39)
{
"harness": "Claude Code (Ralph-Loop mechanism)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "12 hours",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: posttrain-bench not yet in data/benchmarks.json. Prose value 0.37; results table image prints the same result in percent form as 37.1. Methodology: Claude Code with Ralph-Loop for 12 hours, 4 base models across 5 non-LLM-judge benchmarks (AIME2025, BFCL, GPQA Main, GSM8K, HumanEval). Task: autonomous data synthesis, training, evaluation, iteration in 12h.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Methodology · row: OSWorld-Verified · quote_snippet: Increasing Max Steps from 100 to 200 improved task completion rate from 68.70% to 70.06%
{
"harness": "official OSWorld-Verified codebase (testing script announced open-source)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "max steps 200",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Methodology prose states both 68.70% (100 steps) and 70.06% (200 steps); the results table image prints 70.06 for the final configuration. Hero chart image vision-read suggested 75.2 for this row, which conflicts with both text sources; treat 70.06 as the cross-confirmed value and flag the chart read for human review. Protocol: 361 samples from nogdrive collection, relative coordinates 0-1000, resolution 1920x1080, Max Steps 200.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Methodology · row: Video-MME (512-frame condition) · quote_snippet: MiniMax M3 scored 84.6 at 512 frames
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "M3: max output 16K; external: max output 64K",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "LLM-as-a-Judge"
}new-benchmark: video-mme not yet in data/benchmarks.json. Methodology note: official protocol allows 1024 frames but external API limited to 640; M3 scored 84.6 at 512 frames. Separate from the table-image 'VideoMME (w/ sub)' row (85.4, vision-read) - different frame budget and subtitle condition; do not merge. Protocol: 1 FPS, single-frame long edge 336-672px, subtitles interleaved every 30s, official prompt, LLM-as-a-Judge, M3 max output 16K, temperature 1.0, top_p 0.95.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: SWE-Bench Verified · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Claude Code (default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 4,
"aggregation": "mean of 4 runs",
"judge": null
}Vision-read value (unconfirmed) in Coding group: M3=80.5; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: internal infrastructure, Claude Code scaffolding, default system prompt overridden, each test run 4 times and averaged. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):79.9 / 87.6 / 82.9 / 80.6 / 79.6 / 80.6 / - / 80.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: SWE Atlas-QnA · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Mini-SWE-Agent",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "3 hours",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swe-atlas-codebase-qna not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Coding group: M3=37.9; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: sandbox 4C8G, 3h timeout; Claude Sonnet 4.6 / GPT-5.5 / Gemini 3.1 Pro cells from labs.scale.com; Opus 4.7, M2.7 and M3 used Mini-SWE-Agent with evaluation logic aligned to official method - competitor cells are not same-scaffolding. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):11.29 / 45.16 / 45.43 / 13.5 / 31.20 / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: nl2repo · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Claude Code (scaffolding; GPT-5.5 cell used Codex)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "4 hours",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: nl2repo not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Coding group: M3=42.13; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: sandbox 1C2G, 4h timeout; DeepSeek-V4-pro / Kimi-k2.6 / GLM-5.1 cells cited from https://qwen.ai/blog?id=qwen3.7; Claude Code scaffolding for Opus 4.7 / M2.7 / M3 / Gemini 3.1 Pro, GPT-5.5 on Codex. Anti-hack modifications: prompts forbid external info via git clone / pip install; scaffold-monitored Bash analyzed for cheating, cheating commands intercepted with warning. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):34.99 / 56.28 / 52.9 / 21.62 / - / 35.5 / 41 / 42.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: SWE Atlas-Test Writing · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Claude Code (scaffolding)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "3 hours",
"run_count": 4,
"aggregation": "mean of 4 runs",
"judge": null
}new-benchmark: swe-atlas-test-writing not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Coding group: M3=30.83; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.1 Pro cells from labs.scale.com; Opus 4.7 / M2.7 / M3 on internal infra, Claude Code, sandbox 4C8G, 3h timeout, 4 runs averaged. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):18.89 / 38.21 / 42.59 / 29.84 / 31.76 / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: LiveSQLBench · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Claude Code (scaffolding, task prompt overrides default system prompt)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "25 minutes",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: livesqlbench not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Coding group: M3=40.17; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: open-source LiveSQLBench-Base-Full v1 (600 questions / 22 PostgreSQL databases), official workflow, dedicated sandbox per question with pre-installed PostgreSQL, 25-minute timeout, task prompt overrides default system prompt. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):33.17 / 41.00 / 40.17 / 39.83 / - / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: CL-bench · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cl-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Coding group: M3=20.48; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: open-source CL-bench data and rubrics, setup fully aligned with official procedure. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):15.38 / 22.92 / 25.38 / 21.06 / - / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: VIBE-V2 · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Claude Code (scaffolding)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "mean of 3 runs",
"judge": "Agent-as-a-Verifier automated verification"
}new-benchmark: vibe-v2 not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Coding group: M3=50.12; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: internal benchmark (pure front-end and full-stack Web/Android/iOS, build-from-scratch), Claude Code scaffolding, Agent-as-a-Verifier paradigm for program interaction and visual output, unified pipeline with requirement set + containerized deployment + dynamic interaction environment, 3 runs averaged. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):37.89 / 55.87 / 50.50 / 28.00 / - / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: SVG-Bench · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Claude Code (scaffolding)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "mean of 3 runs",
"judge": "VLM rendering-accuracy verification"
}new-benchmark: svg-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Coding group: M3=63.7; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: internal benchmark, inputs text and images, tasks build-from-scratch or edit, Claude Code scaffolding, VLM verifies rendering accuracy, 3 runs averaged. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):48.0 / 62.3 / 58.2 / 59.2 / - / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: PaperBench · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Claude Code (Ralph-Loop mechanism)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "12 hours",
"run_count": null,
"aggregation": null,
"judge": "Opus-4.6"
}new-benchmark: paperbench not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Coding group: M3=52.6; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: Claude Code with Ralph-Loop for 12 hours, 19 papers reproducible without external API, official open-source human expert rubrics, scoring model Opus-4.6. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):30.6 / 58.5 / 57.5 / 46.7 / - / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: BrowseComp · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "WebExplorer agent framework (Liu et al., 2025)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "history discarded once token usage exceeds 64K",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Cowork group: M3=83.52; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: same agent framework as WebExplorer (Liu et al., 2025); all history discarded once token usage exceeds 64K. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):76.3 / 79.3 / 84.4 / 85.9 / 74.7 / 83.4 / 79.3 / 83.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: DRACO · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "internal scaffolding (Deep Research Skill in MiniMax Code)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Claude Opus 4.6"
}new-benchmark: draco not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Cowork group: M3=73.23; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: internal scaffolding (accessible via Deep Research Skill in MiniMax Code); scoring per official rubrics, final score = average across questions; scoring model Claude Opus 4.6; Claude Opus 4.7 cell taken from the Opus 4.7 model card. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):66.77 / 77.7 / - / - / 75.8 / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: GDPval rubrics · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "environment aligned with GDPval-AA scaffolding",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "pointwise scoring on public rubrics"
}new-benchmark: gdpval-rubrics not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Cowork group: M3=74.78; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: internal evaluation using cases from the public GDPval dataset, pointwise scoring on public rubrics, environment aligned with GDPval-AA scaffolding. Distinct protocol from gdpval (leaderboard) and gdpval-aa (Artificial Analysis) rows used by other releases. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):66.44 / 79.8 / 80.66 / 57.82 / 75.65 / 70.32 / 68.26 / 65.12。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: Banker ToolBench · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Claude Code (scaffolding; GPT-5.5 cell used Codex)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "MiniMax M2.7"
}new-benchmark: bankertoolbench not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Cowork group: M3=76.12; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: public BankerToolBench dataset; all models except GPT-5.5 on Claude Code scaffolding (GPT-5.5 on Codex); dataset rubrics scoring; scoring model MiniMax M2.7. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):63.89 / 81.34 / 70.04 / 67.03 / - / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: OfficeQA Pro · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Claude Code (scaffolding; files provided as a file system)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "exact-match scoring"
}new-benchmark: officeqa-pro not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Cowork group: M3=45.1; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: relevant files provided as a file system to simulate realistic scenarios, Claude Code scaffolding, exact-match scoring. Same benchmark id as kimi-k3 release rows. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):- / 43.6 / 52.6 / 18.1 / - / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: SpreadSheetBench-v1 · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "Claude Code (scaffolding)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: spreadsheetbench not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Cowork group: M3=89.35; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: public dataset, Claude Code scaffolding. Chart label carries -v1 suffix. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):84.92 / 88.49 / 88.11 / 56.06 / - / 84.9 / 85.2 / 84.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: YC-Bench · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "official YC-Bench codebase and configuration",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: yc-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Cowork group: M3=2.10M; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: official YC-Bench codebase and configuration, environment aligned with official setup; metric is final assets (fund), so unit is USD not percent. Vision-read cells: M2.7 0, Opus 4.7 2.19M, GPT-5.5 1.28M, Gemini 3.1 Pro 1.05M. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):0 / 2.19M / 1.28M / 1.05M / - / - / - / -。 2026-09-01 audit: value upgraded from not_extracted to reported (cell reads 2.10M = 2,100,000 USD final assets); unit corrected percent -> usd to match metric usd_fund.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: LOCA-Bench (256k) · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "official LOCA-Bench codebase, official react mode",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "Environment Description Length = 256k",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: loca-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Cowork group: M3=49.3; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: official LOCA-Bench codebase, official react mode, Environment Description Length 256k. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):0 / 57 / - / - / - / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: Apex-Agents · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "archipelago codebase, ReAct Toolbelt framework",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Claude Sonnet 4.6"
}new-benchmark: apex-agents not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Cowork group: M3=27.7; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: archipelago codebase, ReAct Toolbelt framework, scoring model Claude Sonnet 4.6. Same benchmark id as kimi-k3 release rows. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):5.6 / 37.2 / 41.7 / 33.4 / 26.2 / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: Claw-Eval · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "official Claw-Eval codebase, General Task Group (161 tasks)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Gemini 3.0 Flash"
}new-benchmark: claw-eval not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Cowork group: M3=74.5; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: official Claw-Eval codebase, General Task Group (161 tasks), scoring model Gemini 3.0 Flash aligned with official model; metric Pass^3. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):49.7 / 71.6 / - / 57.8 / 68.3 / 58.4 / 62.7 / 61.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: OmniDocBench · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "image long edge max 3584 pixels",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: omnidocbench not yet in data/benchmarks.json. Vision-read value (unconfirmed) in MultiModal group: M3=91.6; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: public OmniDocBench v1.5 dataset, official evaluation logic, image long edge max 3584 pixels, reasonable formatting constraints added on top of official prompts; Gemini 3.1 Pro / GPT-5.5 / Opus 4.7 on default API parameters. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):- / 89.3 / 87.5 / 88.1 / 86.9 / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: MMMU-Pro · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed) in MultiModal group: M3=78.1; table cells for other models are vision-read competitor values, see notes column mapping. Maps to existing benchmark mmmu, variant Pro. Methodology: aligned with official evaluation, prompt enforces a format constraint on the model's last line for parsing. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):- / 77 / 81.2 / 80.5 / 74.5 / - / - / 79.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: Video-MMMU · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "highest thinking mode (external models)",
"temperature": 1,
"top_p": 0.95,
"token_budget": "M3: max output 32K; external: max output 64K",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "LLM-as-a-Judge"
}new-benchmark: video-mmmu not yet in data/benchmarks.json. Vision-read value (unconfirmed) in MultiModal group: M3=84.6; table cells for other models are vision-read competitor values, see notes column mapping. Methodology: 1 FPS, max 512 frames, single-frame long edge 672-1008px, official VideoMMMU prompt, LLM-as-a-Judge; M3 max output 32K / temp 1.0 / top_p 0.95; external models max output 64K / temp 0.7 / top_p 0.95 / highest thinking mode - competitor cells not same-configuration. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):- / 83 / 86.4 / 87.9 / - / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: VideoMME (w/ sub) · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "M3: max output 16K; external: max output 64K",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "LLM-as-a-Judge"
}new-benchmark: video-mme not yet in data/benchmarks.json. Vision-read value (unconfirmed) in MultiModal group: M3=85.4; table cells for other models are vision-read competitor values, see notes column mapping. Table row label 'VideoMME (w/ sub)'. Methodology: 1 FPS, max 1024 frames, single-frame long edge 336-672px, subtitles inserted every 30 seconds interleaved into frames, official prompt, LLM-as-a-Judge; M3 max output 16K / temp 1.0 / top_p 0.95. Note: methodology footnote records Claude Opus 4.7 API error rate >20%, results not reported. Keep separate from the prose-verified 512-frame 84.6 row. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):- / - / 89.4 / 87.9 / - / - / - / -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: IMO 2025 · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "MathArena-aligned",
"tools": null,
"shots": null,
"reasoning_effort": "test-time-scaling framework up to 10 iterations",
"temperature": 1,
"top_p": null,
"token_budget": "512k max output tokens",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "dual strong-model grading (GPT-5.4 high effort + Gemini 3.1 Pro high effort), minimum of two judges"
}new-benchmark: imo-2025 not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Reasoning group: M3=35; table cells for other models are vision-read competitor values, see notes column mapping. Table cell '35 / 42' (max score 42, 6 problems). Only M3 has a value in this row; all competitor cells are dashes. Methodology: aligned with MathArena official evaluation, solution normalization then dual strong-model grading on human expert rubrics (GPT-5.4 high effort + Gemini 3.1 Pro high effort), minimum of the two judges; M3 with 512k max output tokens, temperature 1.0, test-time-scaling up to 10 iterations; closed-source baselines are avg@k. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):全列 '-'(M3 单独报告)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark results table image (between the API section and the 'Evaluation Methodology' section) · row: USAMO 2026 · figure: images/19.jpg (archive of filecdn.minimax.chat img_v3_02128_...8686.jpg 全表)
{
"harness": "MathArena-aligned",
"tools": null,
"shots": null,
"reasoning_effort": "test-time-scaling framework up to 10 iterations",
"temperature": 1,
"top_p": null,
"token_budget": "512k max output tokens",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "dual strong-model grading (GPT-5.4 high effort + Gemini 3.1 Pro high effort), minimum of two judges"
}new-benchmark: usamo-2026 not yet in data/benchmarks.json. Vision-read value (unconfirmed) in Reasoning group: M3=36; table cells for other models are vision-read competitor values, see notes column mapping. Table cell '36 / 42' (max score 42). Competitor cells printed in percent (Opus 4.7 52.8%, GPT-5.5 98.21%, Gemini 3.1 Pro 74.40%) - mixed units within one row, do not compare numerically until normalized. Same MathArena-aligned dual-judge protocol as IMO 2025. 视觉转写自归档图 images/19.jpg(官方全表,2026-09-01 复核,与先前读数一致)。同表竞品列(M2.7 / Opus4.7 / GPT5.5 / Gemini3.1Pro / Sonnet4.6 / DSV4Pro / GLM5.1T / K2.6T):- / 52.8% / 98.21% / 74.40% / - / - / - / -。
MiniMax-M2.7
MiniMax-M2.7 是发布结果表中的前代对照列,仅作为 M3 各组基准的比较基线,本次发布未为其报告新评测数值。
- 输入模态
- 官方资料未说明
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。
未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Real-World Tasks / Letting M3 Train Models · row: PostTrainBench (Claude Opus 4.7 column) · quote_snippet: slightly below Opus 4.7 (0.42)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: posttrain-bench not yet in data/benchmarks.json. Competitor score printed in prose; MiniMax methodology notes Opus 4.7 results for DRACO are taken from the Opus 4.7 model card, and for the table image the Opus 4.7 PostTrainBench cell is 42.4 (percent form). Verify against the Anthropic model card before any cross-vendor comparison.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Real-World Tasks / Letting M3 Train Models · row: PostTrainBench (GPT-5.5 column) · quote_snippet: GPT-5.5 (0.39)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: posttrain-bench not yet in data/benchmarks.json. Competitor score printed in prose; results table image prints 39.3 (percent form) for the same cell. Verify against OpenAI source before use.