← 模型目录

MiniMax-M2.7

MiniMax · 2026-03-18 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

MiniMax-M2.7

MiniMax 发布 M2.7「Early Echoes of Self-Evolution」,定位为首个深度参与自身进化的 M 系列模型,支持 Agent Teams、复杂 Skills 与动态工具搜索。评测集中在智能体与编码,亮点如 SWE-Bench Pro 56.2%、MLE-Bench Lite 66.6%。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

swebench-pro 56.22% 模型 minimax-m2-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional Software Engineering / Professional Work · row: SWE-Pro · quote_snippet: On the SWE-Pro benchmark, M2.7 scored 56.22%, nearly approaching Opus's best level

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

swebench-pro id already introduced by prior batches. Later prose: 'matching GPT-5.3-Codex'.

打开官方来源

swebench-multilingual 76.5 模型 minimax-m2-7 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional Software Engineering / Professional Work · row: SWE Multilingual · quote_snippet: such as SWE Multilingual (76.5)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-multilingual already introduced by prior batches.

打开官方来源

multi-swe-bench 52.7 模型 minimax-m2-7 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional Software Engineering / Professional Work · row: Multi SWE Bench · quote_snippet: Multi SWE Bench (52.7)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: multi-swe-bench already introduced by prior batches.

打开官方来源

vibe-pro 55.6% 模型 minimax-m2-7 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional Software Engineering / Professional Work · row: VIBE-Pro · quote_snippet: On the repo-level code generation benchmark VIBE-Pro, M2.7 scored 55.6%, nearly on par with Opus 4.6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: vibe-pro not yet in data/benchmarks/ (MiniMax-upgraded Pro version of the VIBE repo-level generation benchmark; also referenced in minimax-m2-5.json).

打开官方来源

terminalbench 57.0% 模型 minimax-m2-7 · 版本 2 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional Software Engineering / Professional Work · row: Terminal Bench 2 · quote_snippet: deep understanding of complex engineering systems on Terminal Bench 2 (57.0%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Version printed as 'Terminal Bench 2'; harness undisclosed.

打开官方来源

nl2repo 39.8 模型 minimax-m2-7 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional Software Engineering / Professional Work · row: NL2Repo · quote_snippet: On Terminal Bench 2 (57.0%) and NL2Repo (39.8)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

nl2repo id already introduced by prior batches.

打开官方来源

gdpval-aa 1495 模型 minimax-m2-7 · 版本 未说明 · 指标 elo_rating · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional Software Engineering / Professional Work · row: GDPval-AA · figure: images/20.jpg (GDPval-AA panel) · quote_snippet: Its ELO score on GDPval-AA is 1495, the highest among open-source models

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

gdpval-aa id already introduced by prior batches. Elo unit - never aggregate with percent rows. Prose: 'among 45 models, second only to Opus 4.6, Sonnet 4.6, and GPT5.4, surpassing GPT5.3'. Chart panel (vision, 2026-09-01) prints the same comparison in percent-style bars: M2.7 50 / M2.5 35 / Gemini 3.1 Pro 41 / Sonnet 4.6 57 / Opus 4.6 55 / GPT 5.4 58 - ranking matches the prose (above M2.7: Opus 4.6, Sonnet 4.6, GPT-5.4). ELO 1495 from prose stays the primary record.

打开官方来源

toolathlon 46.3% 模型 minimax-m2-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional Software Engineering / Professional Work · row: Toolathon (page spelling) = Tool-Decathlon · quote_snippet: On Toolathon, M2.7 achieved an accuracy of 46.3%, reaching the global top tier

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: tool-decathlon. Page spells it 'Toolathon'; the 46.3 value matches GLM-5.1's Tool-Decathlon column for M2.7 exactly, so the tool-decathlon mapping is near-certain (cross-vendor confirmation) - spelling discrepancy recorded.

打开官方来源

mlebench 66.6% 模型 minimax-m2-7 · 版本 Lite, 22 competitions · 指标 medal_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional Software Engineering / Professional Work · row: MLE-Bench Lite medal rate · quote_snippet: The average medal rate across the three runs was 66.6%, a result second only to Opus-4.6 (75.7%) and GPT-5.4 (71.2%)

{
  "harness": null,
  "tools": "single A30 GPU per competition",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "3 trials x 24 hours iterative evolution each",
  "run_count": 3,
  "aggregation": "average medal rate over 3 runs",
  "judge": null
}

mlebench id exists in data/benchmarks/. Metric is medal rate (medals/competitions), not accuracy. Best single run: 9 gold / 5 silver / 1 bronze of 22. Harness: short-term memory + self-feedback + self-optimization loop (vendor-designed).

打开官方来源

mm-claw 97% 模型 minimax-m2-7 · 版本 40 complex skills · 指标 skill_adherence_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional Software Engineering / Professional Work · row: MM Claw skill compliance · quote_snippet: M2.7 maintained a 97% skill compliance rate across 40 complex skills (each exceeding 2,000 tokens)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mm-claw not yet in data/benchmarks/ (MiniMax internal skill-adherence suite). Metric is a compliance rate over long skills - distinct construct from task accuracy.

打开官方来源

mm-claw 62.7% 模型 minimax-m2-7 · 版本 MM-ClawBench · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: OpenClaw / MM Claw · row: MM-ClawBench · figure: images/20.jpg (MM-ClawBench panel) · quote_snippet: M2.7 achieved a level close to Sonnet 4.6 on this test, with an accuracy of 62.7%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

MM Claw accuracy claim ("level close to Sonnet 4.6"). Chart read (vision, 2026-09-01): M2.7 62.7 vs M2.5 57.6 / Gemini 3.1 Pro 61.8 / Sonnet 4.6 64.2 / Opus 4.6 75.4 / GPT 5.4 73.6. Distinct from the 97% skill-adherence entry (40 complex skills).

打开官方来源

aa-intelligence-index 50 模型 minimax-m2-7 · 版本 Artificial Analysis panel (chart) · 指标 intelligence_index · 单位 points 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Closing benchmark panel · row: Artificial Analysis · figure: images/20.jpg (Artificial Analysis panel)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Artificial Analysis aggregate (integer scale in chart)",
  "judge": null
}

Chart-only value (vision read, 2026-09-01): M2.7 50 vs M2.5 42 / Gemini 3.1 Pro 57 / Sonnet 4.6 52 / Opus 4.6 53 / GPT 5.4 57. No prose claim about AA on the M2.7 page.

打开官方来源