MiniMax-M2.7
MiniMax · 2026-03-18 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
MiniMax-M2.7
MiniMax 发布 M2.7「Early Echoes of Self-Evolution」,定位为首个深度参与自身进化的 M 系列模型,支持 Agent Teams、复杂 Skills 与动态工具搜索。评测集中在智能体与编码,亮点如 SWE-Bench Pro 56.2%、MLE-Bench Lite 66.6%。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional Software Engineering / Professional Work · row: SWE-Pro · quote_snippet: On the SWE-Pro benchmark, M2.7 scored 56.22%, nearly approaching Opus's best level
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}swebench-pro id already introduced by prior batches. Later prose: 'matching GPT-5.3-Codex'.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional Software Engineering / Professional Work · row: SWE Multilingual · quote_snippet: such as SWE Multilingual (76.5)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-multilingual already introduced by prior batches.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional Software Engineering / Professional Work · row: Multi SWE Bench · quote_snippet: Multi SWE Bench (52.7)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: multi-swe-bench already introduced by prior batches.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional Software Engineering / Professional Work · row: VIBE-Pro · quote_snippet: On the repo-level code generation benchmark VIBE-Pro, M2.7 scored 55.6%, nearly on par with Opus 4.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: vibe-pro not yet in data/benchmarks/ (MiniMax-upgraded Pro version of the VIBE repo-level generation benchmark; also referenced in minimax-m2-5.json).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional Software Engineering / Professional Work · row: Terminal Bench 2 · quote_snippet: deep understanding of complex engineering systems on Terminal Bench 2 (57.0%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Version printed as 'Terminal Bench 2'; harness undisclosed.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional Software Engineering / Professional Work · row: NL2Repo · quote_snippet: On Terminal Bench 2 (57.0%) and NL2Repo (39.8)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}nl2repo id already introduced by prior batches.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional Software Engineering / Professional Work · row: GDPval-AA · figure: images/20.jpg (GDPval-AA panel) · quote_snippet: Its ELO score on GDPval-AA is 1495, the highest among open-source models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}gdpval-aa id already introduced by prior batches. Elo unit - never aggregate with percent rows. Prose: 'among 45 models, second only to Opus 4.6, Sonnet 4.6, and GPT5.4, surpassing GPT5.3'. Chart panel (vision, 2026-09-01) prints the same comparison in percent-style bars: M2.7 50 / M2.5 35 / Gemini 3.1 Pro 41 / Sonnet 4.6 57 / Opus 4.6 55 / GPT 5.4 58 - ranking matches the prose (above M2.7: Opus 4.6, Sonnet 4.6, GPT-5.4). ELO 1495 from prose stays the primary record.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional Software Engineering / Professional Work · row: Toolathon (page spelling) = Tool-Decathlon · quote_snippet: On Toolathon, M2.7 achieved an accuracy of 46.3%, reaching the global top tier
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tool-decathlon. Page spells it 'Toolathon'; the 46.3 value matches GLM-5.1's Tool-Decathlon column for M2.7 exactly, so the tool-decathlon mapping is near-certain (cross-vendor confirmation) - spelling discrepancy recorded.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional Software Engineering / Professional Work · row: MLE-Bench Lite medal rate · quote_snippet: The average medal rate across the three runs was 66.6%, a result second only to Opus-4.6 (75.7%) and GPT-5.4 (71.2%)
{
"harness": null,
"tools": "single A30 GPU per competition",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "3 trials x 24 hours iterative evolution each",
"run_count": 3,
"aggregation": "average medal rate over 3 runs",
"judge": null
}mlebench id exists in data/benchmarks/. Metric is medal rate (medals/competitions), not accuracy. Best single run: 9 gold / 5 silver / 1 bronze of 22. Harness: short-term memory + self-feedback + self-optimization loop (vendor-designed).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional Software Engineering / Professional Work · row: MM Claw skill compliance · quote_snippet: M2.7 maintained a 97% skill compliance rate across 40 complex skills (each exceeding 2,000 tokens)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mm-claw not yet in data/benchmarks/ (MiniMax internal skill-adherence suite). Metric is a compliance rate over long skills - distinct construct from task accuracy.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: OpenClaw / MM Claw · row: MM-ClawBench · figure: images/20.jpg (MM-ClawBench panel) · quote_snippet: M2.7 achieved a level close to Sonnet 4.6 on this test, with an accuracy of 62.7%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}MM Claw accuracy claim ("level close to Sonnet 4.6"). Chart read (vision, 2026-09-01): M2.7 62.7 vs M2.5 57.6 / Gemini 3.1 Pro 61.8 / Sonnet 4.6 64.2 / Opus 4.6 75.4 / GPT 5.4 73.6. Distinct from the 97% skill-adherence entry (40 complex skills).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Closing benchmark panel · row: Artificial Analysis · figure: images/20.jpg (Artificial Analysis panel)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Artificial Analysis aggregate (integer scale in chart)",
"judge": null
}Chart-only value (vision read, 2026-09-01): M2.7 50 vs M2.5 42 / Gemini 3.1 Pro 57 / Sonnet 4.6 52 / Opus 4.6 53 / GPT 5.4 57. No prose claim about AA on the M2.7 page.