MiniMax-M2.5
MiniMax · 2026-02-12 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
MiniMax-M2.5
MiniMax 将 M2.5 定位为「为真实生产力打造」,披露 20 万+ 真实环境 RL 训练与 Forge 智能体原生框架,并随附同能力 2 倍速的 M2.5-Lightning 档。已收录 22 项评测集中于编码、搜索与 Agent 工具调用:亮点 SWE-bench Verified 80.2%、tau2-Bench Telecom 97.8%。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 0.15 / 输出 1.2 · M2.5 为 50 TPS 档、价格为 M2.5-Lightning(100 TPS,$0.3/$2.4 每百万 tokens)的一半
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Search and Tool calling / Office work · row: SWE-Bench Verified · quote_snippet: boasting scores of 80.2% in SWE-Bench Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "3.52M tokens per task average",
"turn_limit": null,
"time_limit": "22.8 minutes end-to-end average per task",
"run_count": null,
"aggregation": null,
"judge": null
}Token and time economics disclosed alongside the score - rare. Harness not named.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Search and Tool calling / Office work · row: SWE-Bench Verified on Droid · quote_snippet: On Droid: 79.7(M2.5) > 78.9(Opus 4.6)
{
"harness": "Droid",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Cross-harness generalization disclosure: same benchmark, different agent harness. Opus 4.6 78.9 printed as comparison.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Search and Tool calling / Office work · row: SWE-Bench Verified on OpenCode · quote_snippet: On OpenCode: 76.1(M2.5) > 75.9(Opus 4.6)
{
"harness": "OpenCode",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Second harness row; three same-benchmark different-harness scores (80.2 default / 79.7 Droid / 76.1 OpenCode) demonstrate harness sensitivity ~4 points.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Search and Tool calling / Office work · row: Multi-SWE-Bench · quote_snippet: 51.3% in Multi-SWE-Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: multi-swe-bench already introduced by prior batches.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Search and Tool calling / Office work · row: BrowseComp (with context management) · quote_snippet: 76.3% in BrowseComp (with context management)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp already introduced by prior batches. Condition explicitly 'with context management' - not comparable to no-CM rows in GLM files.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Search and Tool calling / Office work · row: GDPval-MM (internal Cowork Agent evaluation) · quote_snippet: it achieved an average win rate of 59.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "token costs monitored across the workflow",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "pairwise comparison of deliverable quality and trajectory professionalism"
}new-benchmark: gdpval-mm (MiniMax internal multimodal GDPval-style Cowork Agent evaluation) not yet in data/benchmarks/. Win rate vs mainstream models under pairwise comparison, not an absolute accuracy.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: VIBE-Pro (AVG) · figure: images/20.webp + images/28.webp (M2.5 red bar 54.2; subset panel images/44.webp)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: vibe-pro (vendor-upgraded Pro version of VIBE). Value read from appendix chart (vision, 2026-09-01): M2.5 54.2 AVG vs M2.1 42.4 / Opus 4.5 55.2 / Opus 4.6 55.6 / Gemini 3 Pro 36.9. Subset panel (images/44.webp): Web 36.9, Simulation 81.4, Android 50.6, iOS 47.9 - mean of 4 subsets equals 54.2 AVG. Prose itself only claims parity with Opus 4.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Search and Tool calling · row: RISE · figure: images/36.webp (M2.5 red bar 50.2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: rise (MiniMax-built realistic interactive search evaluation over professional human-expert tasks) not yet in data/benchmarks/. Value read from chart (vision, 2026-09-01): M2.5 50.2 vs M2.1 34 / Opus 4.5 50.5 / Opus 4.6 62.5 / Gemini 3 Pro 36.8 / GPT-5.2 50. Round-efficiency claim: ~20% fewer rounds than M2.1 across BrowseComp/Wide Search/RISE.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: SWE-Bench Pro · figure: images/20.webp + images/28.webp
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}Chart row read (vision, 2026-09-01): M2.5 55.4 vs M2.1 49.7 / Opus 4.5 56.9 / Opus 4.6 55.4 / Gemini 3 Pro 54.1 / GPT-5.2 55.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: SWE-Bench Multilingual · figure: images/28.webp
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}Chart row read (vision, 2026-09-01): M2.5 74.1 vs M2.1 71.9 / Opus 4.5 77.5 / Opus 4.6 77.8 / Gemini 3 Pro 65 / GPT-5.2 72.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: Terminal Bench 2 · figure: images/28.webp
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}Terminal Bench 2 protocol from appendix: Claude Code 2.0.64 scaffolding, modified Dockerfiles, 4 runs averaged. Chart read (vision, 2026-09-01): M2.5 51.7 vs M2.1 47.9 / Opus 4.5 53.4 / Opus 4.6 55.1 / Gemini 3 Pro 54 / GPT-5.2 54.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: BFCL multi-turn · figure: images/20.webp + images/36.webp
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}Chart row read (vision, 2026-09-01): M2.5 76.8 vs M2.1 37.4 / Opus 4.5 68 / Opus 4.6 63.3 / Gemini 3 Pro 61.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: tau2 Telecom · figure: images/36.webp
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}Chart row read (vision, 2026-09-01): M2.5 97.8 vs M2.1 87 / Opus 4.5 98.2 / Opus 4.6 99.3 / Gemini 3 Pro 98 / GPT-5.2 98.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: Wide Search · figure: images/36.webp
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}Chart row read (vision, 2026-09-01): M2.5 70.3 vs M2.1 63.2 / Opus 4.5 76.2 / Opus 4.6 79.4 / Gemini 3 Pro 57.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: MEWC · figure: images/20.webp + images/52.webp
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}new-benchmark: mewc (MEWC, Microsoft Excel World Championship; 179 problems from main + regional divisions 2021-2026, MiniMax internal) not yet in data/benchmarks/. Chart read (vision, 2026-09-01): M2.5 74.4 vs M2.1 55.6 / Opus 4.5 82.1 / Opus 4.6 89.8 / Gemini 3 Pro 78.7 / GPT-5.2 41.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: Finance Modeling · figure: images/52.webp
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}new-benchmark: finance-modeling (MiniMax internal, expert-built financial modeling rubric, 3 runs averaged) not yet in data/benchmarks/. Chart read (vision, 2026-09-01): M2.5 21.6 vs M2.1 17.3 / Opus 4.5 30.1 / Opus 4.6 33.2 / Gemini 3 Pro 15 / GPT-5.2 20.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: AIME25 · figure: images/84.webp (AA Intelligence Index sub-table)
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}AIME25 row: M2.5 86.3 vs M2.1 83.0 / Sonnet 4.5 88.0 / Opus 4.5 91.0 / Opus 4.6 95.6 / Gemini 3 Pro 96.0 / GPT-5.2 98.0. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: GPQA-D · figure: images/84.webp (AA Intelligence Index sub-table)
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}GPQA-D row: M2.5 85.2 vs M2.1 83.0 / Sonnet 4.5 83.0 / Opus 4.5 87.0 / Opus 4.6 90.0 / Gemini 3 Pro 91.0 / GPT-5.2 90.0. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: HLE w/o tools · figure: images/84.webp (AA Intelligence Index sub-table)
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}HLE w/o tools row: M2.5 19.4 vs M2.1 22.2 / Sonnet 4.5 17.3 / Opus 4.5 28.4 / Opus 4.6 30.7 / Gemini 3 Pro 37.2 / GPT-5.2 31.4. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: SciCode · figure: images/84.webp (AA Intelligence Index sub-table)
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}SciCode row: M2.5 44.4 vs M2.1 41.0 / Sonnet 4.5 45.0 / Opus 4.5 50.0 / Opus 4.6 52.0 / Gemini 3 Pro 56.0 / GPT-5.2 52.0. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: IFBench · figure: images/84.webp (AA Intelligence Index sub-table)
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}IFBench row: M2.5 70.0 vs M2.1 70.0 / Sonnet 4.5 57.0 / Opus 4.5 58.0 / Opus 4.6 53.0 / Gemini 3 Pro 70.0 / GPT-5.2 75.0. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding / Appendix further benchmark results · row: AA-LCR · figure: images/84.webp (AA Intelligence Index sub-table)
{
"harness": "Claude Code (internal infrastructure, default system prompt overridden)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "4 runs averaged",
"aggregation": null,
"judge": null
}AA-LCR row: M2.5 69.5 vs M2.1 62.0 / Sonnet 4.5 66.0 / Opus 4.5 74.0 / Opus 4.6 71.0 / Gemini 3 Pro 71.0 / GPT-5.2 73.0. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).