← 模型目录

MiniMax-M2

MiniMax · 2025-10-27 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

MiniMax-M2

发布文以“简中之巧(Ingenious in Simplicity)”为题,将 M2 定位为 agent 优先的开源权重模型,价格约为 Claude Sonnet 的 8%、速度约 2 倍。9 项评测集中于智能体编码、工具调用与搜索,亮点为 GAIA(纯文本)75.7 与 τ²-Bench 77.2。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 0.3 / 输出 1.2 · 页面同时给出人民币定价:输入 ¥2.1、输出 ¥8.4 每百万 tokens;另载限时免费试用至 2025-11-07

本变体的评测证据

aa-intelligence-index 61 模型 minimax-m2 · 版本 Intelligence Index v3.0 · 指标 intelligence_index · 单位 points 来源等级 A · third_party_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: paragraph after agentic comparison chart ('on the popular Artificial Analysis benchmark') · row: Artificial Analysis · figure: images/13.png (archive of AA Intelligence Index v3.0 chart) · quote_snippet: on the popular Artificial Analysis benchmark, which integrates 10 test tasks, our model ranked in the top five globally

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Artificial Analysis aggregate over 10 evaluations",
  "judge": null
}

aa-intelligence-index id already introduced by prior batches, still not in data/benchmarks/. Prose rank claim (top five globally) verified; index v3.0 aggregates MMLU-Pro, GPQA Diamond, HLE, LiveCodeBench, SciCode, AIME 2025, IFBench, AA-LCR, Terminal-Bench Hard, tau^2-Bench Telecom. Vision-assisted read of the reprinted AA chart (unconfirmed, third-party values): MiniMax-M2 61 rank 5; GPT-5 (high) / GPT-5 Codex (high) 68, Grok 4 65, Claude 4.5 Sonnet 63, GLM-4.6 56, Qwen3 Max 55, DeepSeek V3.2 Exp 57, Kimi K2 0905 50. Attribution is third_party_reported (AA runs the eval); excluded from vendor self-report counts. 视觉转写自归档图 images/13.png(2026-09-01):图中 MiniMax-M2 柱标注 61(该图为整数刻度,无小数);Kimi K2 0905 同图 50。

打开官方来源

swebench 69.4 模型 minimax-m2 · 版本 Verified · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: agentic comparison chart after 'We compared M2 with several mainstream models' · table: large comparison table image (8676x3593 PNG) · row: SWE-bench Verified · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Page prints no protocol footnotes (unlike the M3 post which has an Evaluation Methodology section) - all protocol fields null. Vision-assisted read (unconfirmed): SWE-bench Verified MiniMax-M2 69.4; DeepSeek-V3.2 67.8, GLM-4.6 68.0, Kimi K2 0905 69.2, Gemini 2.5 Pro 63.8, Claude Sonnet 4.5 77.2, GPT-5 (thinking) 74.9. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 67.8, GLM-4.6 68.0, Kimi K2 0905 69.2, Gemini 2.5 Pro 63.8, Claude Sonnet 4.5 77.2, GPT-5 (thinking) 74.9。

打开官方来源

multi-swe-bench 36.2 模型 minimax-m2 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: Multi-SWE-Bench · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: multi-swe-bench already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read (unconfirmed): Multi-SWE-Bench MiniMax-M2 36.2; DeepSeek-V3.2 30.6, GLM-4.6 30.0, Kimi K2 0905 33.5, Claude Sonnet 4.5 44.3; Gemini 2.5 Pro and GPT-5 not reported. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 30.6, GLM-4.6 30.0, Kimi K2 0905 33.5, Claude Sonnet 4.5 44.3(Gemini/OpenAI 该图未列)。

打开官方来源

terminalbench 46.3 模型 minimax-m2 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: Terminal-Bench · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): Terminal-Bench MiniMax-M2 46.3; DeepSeek-V3.2 37.7, GLM-4.6 40.5, Kimi K2 0905 44.5, Gemini 2.5 Pro 25.3, Claude Sonnet 4.5 50.0, GPT-5 (thinking) 43.8. Version/harness not printed; note Kimi K2 0905's Terminus-harness 44.5 matches, suggesting same-harness values, but this is unconfirmed. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 37.7, GLM-4.6 40.5, Kimi K2 0905 44.5, Gemini 2.5 Pro 25.3, Claude Sonnet 4.5 50.0, GPT-5 (thinking) 43.8。

打开官方来源

artifactsbench 66.8 模型 minimax-m2 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: ArtifactsBench · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: artifactsbench not yet in data/benchmarks/. Vision-assisted read (unconfirmed): ArtifactsBench MiniMax-M2 66.8; DeepSeek-V3.2 55.8, GLM-4.6 59.8, Kimi K2 0905 54.2, Gemini 2.5 Pro 57.7, Claude Sonnet 4.5 61.5, GPT-5 73.0. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 55.8, GLM-4.6 59.8, Kimi K2 0905 54.2, Gemini 2.5 Pro 57.7, Claude Sonnet 4.5 61.5, GPT-5 (thinking) 73.0。

打开官方来源

tau-bench 77.2 模型 minimax-m2 · 版本 2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: tau^2-Bench · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): tau^2-Bench MiniMax-M2 77.2; DeepSeek-V3.2 66.7, GLM-4.6 75.9, Kimi K2 0905 70.3, Gemini 2.5 Pro 59.2, Claude Sonnet 4.5 84.7, GPT-5 80.1. Weighted vs per-domain split not printed; GLM-4.6's chart labels its own cell 'weighted'. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 66.7, GLM-4.6 75.9, Kimi K2 0905 70.3, Gemini 2.5 Pro 59.2, Claude Sonnet 4.5 84.7, GPT-5 (thinking) 80.1。

打开官方来源

gaia 75.7 模型 minimax-m2 · 版本 text only · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: GAIA (text only) · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): GAIA (text only) MiniMax-M2 75.7; DeepSeek-V3.2 63.5, GLM-4.6 71.9, Kimi K2 0905 60.2, Gemini 2.5 Pro 60.2, Claude Sonnet 4.5 71.2, GPT-5 76.4. Text-only variant differs from full GAIA. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 63.5, GLM-4.6 71.9, Kimi K2 0905 60.2, Gemini 2.5 Pro 60.2, Claude Sonnet 4.5 71.2, GPT-5 (thinking) 76.4。

打开官方来源

browsecomp 44.0 模型 minimax-m2 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: BrowseComp · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read (unconfirmed): BrowseComp MiniMax-M2 44.0; DeepSeek-V3.2 40.1, GLM-4.6 45.1, Kimi K2 0905 14.1, Gemini 2.5 Pro 9.9, Claude Sonnet 4.5 19.6, GPT-5 54.9. Deep-search scaffold not described on this page. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 40.1, GLM-4.6 45.1, Kimi K2 0905 14.1, Gemini 2.5 Pro 9.9, Claude Sonnet 4.5 19.6, GPT-5 (thinking) 54.9。 2026-09-01 audit: display aligned to chart cell text "44.0" (images/12.png).

打开官方来源

finsearchcomp 65.5 模型 minimax-m2 · 版本 global · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: FinSearchComp-global · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

finsearchcomp id already introduced by prior batches (Kimi row used T3 variant; this page uses the global split - different variants, do not merge). Vision-assisted read (unconfirmed): FinSearchComp-global MiniMax-M2 65.5; DeepSeek-V3.2 26.2, GLM-4.6 29.2, Kimi K2 0905 29.5, Gemini 2.5 Pro 42.6, Claude Sonnet 4.5 60.8, GPT-5 63.9. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 26.2, GLM-4.6 29.2, Kimi K2 0905 29.5, Gemini 2.5 Pro 42.6, Claude Sonnet 4.5 60.8, GPT-5 (thinking) 63.9。

打开官方来源