MiniMax-M2
MiniMax · 2025-10-27 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
MiniMax-M2
发布文以“简中之巧(Ingenious in Simplicity)”为题,将 M2 定位为 agent 优先的开源权重模型,价格约为 Claude Sonnet 的 8%、速度约 2 倍。9 项评测集中于智能体编码、工具调用与搜索,亮点为 GAIA(纯文本)75.7 与 τ²-Bench 77.2。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 0.3 / 输出 1.2 · 页面同时给出人民币定价:输入 ¥2.1、输出 ¥8.4 每百万 tokens;另载限时免费试用至 2025-11-07
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: paragraph after agentic comparison chart ('on the popular Artificial Analysis benchmark') · row: Artificial Analysis · figure: images/13.png (archive of AA Intelligence Index v3.0 chart) · quote_snippet: on the popular Artificial Analysis benchmark, which integrates 10 test tasks, our model ranked in the top five globally
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Artificial Analysis aggregate over 10 evaluations",
"judge": null
}aa-intelligence-index id already introduced by prior batches, still not in data/benchmarks/. Prose rank claim (top five globally) verified; index v3.0 aggregates MMLU-Pro, GPQA Diamond, HLE, LiveCodeBench, SciCode, AIME 2025, IFBench, AA-LCR, Terminal-Bench Hard, tau^2-Bench Telecom. Vision-assisted read of the reprinted AA chart (unconfirmed, third-party values): MiniMax-M2 61 rank 5; GPT-5 (high) / GPT-5 Codex (high) 68, Grok 4 65, Claude 4.5 Sonnet 63, GLM-4.6 56, Qwen3 Max 55, DeepSeek V3.2 Exp 57, Kimi K2 0905 50. Attribution is third_party_reported (AA runs the eval); excluded from vendor self-report counts. 视觉转写自归档图 images/13.png(2026-09-01):图中 MiniMax-M2 柱标注 61(该图为整数刻度,无小数);Kimi K2 0905 同图 50。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: agentic comparison chart after 'We compared M2 with several mainstream models' · table: large comparison table image (8676x3593 PNG) · row: SWE-bench Verified · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Page prints no protocol footnotes (unlike the M3 post which has an Evaluation Methodology section) - all protocol fields null. Vision-assisted read (unconfirmed): SWE-bench Verified MiniMax-M2 69.4; DeepSeek-V3.2 67.8, GLM-4.6 68.0, Kimi K2 0905 69.2, Gemini 2.5 Pro 63.8, Claude Sonnet 4.5 77.2, GPT-5 (thinking) 74.9. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 67.8, GLM-4.6 68.0, Kimi K2 0905 69.2, Gemini 2.5 Pro 63.8, Claude Sonnet 4.5 77.2, GPT-5 (thinking) 74.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: Multi-SWE-Bench · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: multi-swe-bench already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read (unconfirmed): Multi-SWE-Bench MiniMax-M2 36.2; DeepSeek-V3.2 30.6, GLM-4.6 30.0, Kimi K2 0905 33.5, Claude Sonnet 4.5 44.3; Gemini 2.5 Pro and GPT-5 not reported. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 30.6, GLM-4.6 30.0, Kimi K2 0905 33.5, Claude Sonnet 4.5 44.3(Gemini/OpenAI 该图未列)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: Terminal-Bench · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): Terminal-Bench MiniMax-M2 46.3; DeepSeek-V3.2 37.7, GLM-4.6 40.5, Kimi K2 0905 44.5, Gemini 2.5 Pro 25.3, Claude Sonnet 4.5 50.0, GPT-5 (thinking) 43.8. Version/harness not printed; note Kimi K2 0905's Terminus-harness 44.5 matches, suggesting same-harness values, but this is unconfirmed. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 37.7, GLM-4.6 40.5, Kimi K2 0905 44.5, Gemini 2.5 Pro 25.3, Claude Sonnet 4.5 50.0, GPT-5 (thinking) 43.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: ArtifactsBench · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: artifactsbench not yet in data/benchmarks/. Vision-assisted read (unconfirmed): ArtifactsBench MiniMax-M2 66.8; DeepSeek-V3.2 55.8, GLM-4.6 59.8, Kimi K2 0905 54.2, Gemini 2.5 Pro 57.7, Claude Sonnet 4.5 61.5, GPT-5 73.0. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 55.8, GLM-4.6 59.8, Kimi K2 0905 54.2, Gemini 2.5 Pro 57.7, Claude Sonnet 4.5 61.5, GPT-5 (thinking) 73.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: tau^2-Bench · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): tau^2-Bench MiniMax-M2 77.2; DeepSeek-V3.2 66.7, GLM-4.6 75.9, Kimi K2 0905 70.3, Gemini 2.5 Pro 59.2, Claude Sonnet 4.5 84.7, GPT-5 80.1. Weighted vs per-domain split not printed; GLM-4.6's chart labels its own cell 'weighted'. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 66.7, GLM-4.6 75.9, Kimi K2 0905 70.3, Gemini 2.5 Pro 59.2, Claude Sonnet 4.5 84.7, GPT-5 (thinking) 80.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: GAIA (text only) · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): GAIA (text only) MiniMax-M2 75.7; DeepSeek-V3.2 63.5, GLM-4.6 71.9, Kimi K2 0905 60.2, Gemini 2.5 Pro 60.2, Claude Sonnet 4.5 71.2, GPT-5 76.4. Text-only variant differs from full GAIA. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 63.5, GLM-4.6 71.9, Kimi K2 0905 60.2, Gemini 2.5 Pro 60.2, Claude Sonnet 4.5 71.2, GPT-5 (thinking) 76.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: BrowseComp · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read (unconfirmed): BrowseComp MiniMax-M2 44.0; DeepSeek-V3.2 40.1, GLM-4.6 45.1, Kimi K2 0905 14.1, Gemini 2.5 Pro 9.9, Claude Sonnet 4.5 19.6, GPT-5 54.9. Deep-search scaffold not described on this page. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 40.1, GLM-4.6 45.1, Kimi K2 0905 14.1, Gemini 2.5 Pro 9.9, Claude Sonnet 4.5 19.6, GPT-5 (thinking) 54.9。 2026-09-01 audit: display aligned to chart cell text "44.0" (images/12.png).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: agentic comparison chart · table: large comparison table image (8676x3593 PNG) · row: FinSearchComp-global · figure: images/12.png (archive of filecdn.minimax.chat 6379df73-...9328c.PNG,分组柱状图面板)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}finsearchcomp id already introduced by prior batches (Kimi row used T3 variant; this page uses the global split - different variants, do not merge). Vision-assisted read (unconfirmed): FinSearchComp-global MiniMax-M2 65.5; DeepSeek-V3.2 26.2, GLM-4.6 29.2, Kimi K2 0905 29.5, Gemini 2.5 Pro 42.6, Claude Sonnet 4.5 60.8, GPT-5 63.9. 视觉转写自归档图 images/12.png(2026-09-01 复核,与先前读数一致)。同图竞品:DeepSeek-V3.2 26.2, GLM-4.6 29.2, Kimi K2 0905 29.5, Gemini 2.5 Pro 42.6, Claude Sonnet 4.5 60.8, GPT-5 (thinking) 63.9。