← 模型目录

MiniMax-M2.5

MiniMax · 2026-02-12 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

MiniMax-M2.5

MiniMax 将 M2.5 定位为「为真实生产力打造」,披露 20 万+ 真实环境 RL 训练与 Forge 智能体原生框架,并随附同能力 2 倍速的 M2.5-Lightning 档。已收录 22 项评测集中于编码、搜索与 Agent 工具调用:亮点 SWE-bench Verified 80.2%、tau2-Bench Telecom 97.8%。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 0.15 / 输出 1.2 · M2.5 为 50 TPS 档、价格为 M2.5-Lightning(100 TPS,$0.3/$2.4 每百万 tokens)的一半

本变体的评测证据

swebench 80.2% 模型 minimax-m2-5 · 版本 Verified · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Search and Tool calling / Office work · row: SWE-Bench Verified · quote_snippet: boasting scores of 80.2% in SWE-Bench Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "3.52M tokens per task average",
  "turn_limit": null,
  "time_limit": "22.8 minutes end-to-end average per task",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Token and time economics disclosed alongside the score - rare. Harness not named.

打开官方来源

swebench 79.7 模型 minimax-m2-5 · 版本 Verified, Droid harness (OOD generalization test) · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Search and Tool calling / Office work · row: SWE-Bench Verified on Droid · quote_snippet: On Droid: 79.7(M2.5) > 78.9(Opus 4.6)

{
  "harness": "Droid",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Cross-harness generalization disclosure: same benchmark, different agent harness. Opus 4.6 78.9 printed as comparison.

打开官方来源

swebench 76.1 模型 minimax-m2-5 · 版本 Verified, OpenCode harness (OOD generalization test) · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Search and Tool calling / Office work · row: SWE-Bench Verified on OpenCode · quote_snippet: On OpenCode: 76.1(M2.5) > 75.9(Opus 4.6)

{
  "harness": "OpenCode",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Second harness row; three same-benchmark different-harness scores (80.2 default / 79.7 Droid / 76.1 OpenCode) demonstrate harness sensitivity ~4 points.

打开官方来源

multi-swe-bench 51.3% 模型 minimax-m2-5 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Search and Tool calling / Office work · row: Multi-SWE-Bench · quote_snippet: 51.3% in Multi-SWE-Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: multi-swe-bench already introduced by prior batches.

打开官方来源

browsecomp 76.3% 模型 minimax-m2-5 · 版本 with context management · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Search and Tool calling / Office work · row: BrowseComp (with context management) · quote_snippet: 76.3% in BrowseComp (with context management)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp already introduced by prior batches. Condition explicitly 'with context management' - not comparable to no-CM rows in GLM files.

打开官方来源

gdpval-mm 59.0% 模型 minimax-m2-5 · 版本 未说明 · 指标 win_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Search and Tool calling / Office work · row: GDPval-MM (internal Cowork Agent evaluation) · quote_snippet: it achieved an average win rate of 59.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "token costs monitored across the workflow",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "pairwise comparison of deliverable quality and trajectory professionalism"
}

new-benchmark: gdpval-mm (MiniMax internal multimodal GDPval-style Cowork Agent evaluation) not yet in data/benchmarks/. Win rate vs mainstream models under pairwise comparison, not an absolute accuracy.

打开官方来源

vibe-pro 54.2 模型 minimax-m2-5 · 版本 未说明 · 指标 avg_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: VIBE-Pro (AVG) · figure: images/20.webp + images/28.webp (M2.5 red bar 54.2; subset panel images/44.webp)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: vibe-pro (vendor-upgraded Pro version of VIBE). Value read from appendix chart (vision, 2026-09-01): M2.5 54.2 AVG vs M2.1 42.4 / Opus 4.5 55.2 / Opus 4.6 55.6 / Gemini 3 Pro 36.9. Subset panel (images/44.webp): Web 36.9, Simulation 81.4, Android 50.6, iOS 47.9 - mean of 4 subsets equals 54.2 AVG. Prose itself only claims parity with Opus 4.5.

打开官方来源

rise 50.2 模型 minimax-m2-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Search and Tool calling · row: RISE · figure: images/36.webp (M2.5 red bar 50.2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: rise (MiniMax-built realistic interactive search evaluation over professional human-expert tasks) not yet in data/benchmarks/. Value read from chart (vision, 2026-09-01): M2.5 50.2 vs M2.1 34 / Opus 4.5 50.5 / Opus 4.6 62.5 / Gemini 3 Pro 36.8 / GPT-5.2 50. Round-efficiency claim: ~20% fewer rounds than M2.1 across BrowseComp/Wide Search/RISE.

打开官方来源

swebench-pro 55.4 模型 minimax-m2-5 · 版本 Pro (2025-09) · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: SWE-Bench Pro · figure: images/20.webp + images/28.webp

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

Chart row read (vision, 2026-09-01): M2.5 55.4 vs M2.1 49.7 / Opus 4.5 56.9 / Opus 4.6 55.4 / Gemini 3 Pro 54.1 / GPT-5.2 55.6.

打开官方来源

swebench-multilingual 74.1 模型 minimax-m2-5 · 版本 v1 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: SWE-Bench Multilingual · figure: images/28.webp

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

Chart row read (vision, 2026-09-01): M2.5 74.1 vs M2.1 71.9 / Opus 4.5 77.5 / Opus 4.6 77.8 / Gemini 3 Pro 65 / GPT-5.2 72.

打开官方来源

terminalbench 51.7 模型 minimax-m2-5 · 版本 2 (Claude Code 2.0.64 scaffolding, sandbox 8-core/16GB, timeout 7200s) · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: Terminal Bench 2 · figure: images/28.webp

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

Terminal Bench 2 protocol from appendix: Claude Code 2.0.64 scaffolding, modified Dockerfiles, 4 runs averaged. Chart read (vision, 2026-09-01): M2.5 51.7 vs M2.1 47.9 / Opus 4.5 53.4 / Opus 4.6 55.1 / Gemini 3 Pro 54 / GPT-5.2 54.

打开官方来源

bfcl 76.8 模型 minimax-m2-5 · 版本 multi-turn · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: BFCL multi-turn · figure: images/20.webp + images/36.webp

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

Chart row read (vision, 2026-09-01): M2.5 76.8 vs M2.1 37.4 / Opus 4.5 68 / Opus 4.6 63.3 / Gemini 3 Pro 61.

打开官方来源

tau2-bench 97.8 模型 minimax-m2-5 · 版本 Telecom · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: tau2 Telecom · figure: images/36.webp

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

Chart row read (vision, 2026-09-01): M2.5 97.8 vs M2.1 87 / Opus 4.5 98.2 / Opus 4.6 99.3 / Gemini 3 Pro 98 / GPT-5.2 98.7.

打开官方来源

widesearch 70.3 模型 minimax-m2-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: Wide Search · figure: images/36.webp

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

Chart row read (vision, 2026-09-01): M2.5 70.3 vs M2.1 63.2 / Opus 4.5 76.2 / Opus 4.6 79.4 / Gemini 3 Pro 57.

打开官方来源

mewc 74.4 模型 minimax-m2-5 · 版本 179 problems, Excel esports 2021-2026 · 指标 score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: MEWC · figure: images/20.webp + images/52.webp

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

new-benchmark: mewc (MEWC, Microsoft Excel World Championship; 179 problems from main + regional divisions 2021-2026, MiniMax internal) not yet in data/benchmarks/. Chart read (vision, 2026-09-01): M2.5 74.4 vs M2.1 55.6 / Opus 4.5 82.1 / Opus 4.6 89.8 / Gemini 3 Pro 78.7 / GPT-5.2 41.3.

打开官方来源

finance-modeling 21.6 模型 minimax-m2-5 · 版本 expert rubric, end-to-end research + Excel · 指标 score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: Finance Modeling · figure: images/52.webp

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

new-benchmark: finance-modeling (MiniMax internal, expert-built financial modeling rubric, 3 runs averaged) not yet in data/benchmarks/. Chart read (vision, 2026-09-01): M2.5 21.6 vs M2.1 17.3 / Opus 4.5 30.1 / Opus 4.6 33.2 / Gemini 3 Pro 15 / GPT-5.2 20.

打开官方来源

aime-25 86.3 模型 minimax-m2-5 · 版本 AIME25 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: AIME25 · figure: images/84.webp (AA Intelligence Index sub-table)

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

AIME25 row: M2.5 86.3 vs M2.1 83.0 / Sonnet 4.5 88.0 / Opus 4.5 91.0 / Opus 4.6 95.6 / Gemini 3 Pro 96.0 / GPT-5.2 98.0. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).

打开官方来源

gpqa 85.2 模型 minimax-m2-5 · 版本 GPQA-D · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: GPQA-D · figure: images/84.webp (AA Intelligence Index sub-table)

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

GPQA-D row: M2.5 85.2 vs M2.1 83.0 / Sonnet 4.5 83.0 / Opus 4.5 87.0 / Opus 4.6 90.0 / Gemini 3 Pro 91.0 / GPT-5.2 90.0. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).

打开官方来源

hlehle 19.4 模型 minimax-m2-5 · 版本 w/o tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: HLE w/o tools · figure: images/84.webp (AA Intelligence Index sub-table)

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

HLE w/o tools row: M2.5 19.4 vs M2.1 22.2 / Sonnet 4.5 17.3 / Opus 4.5 28.4 / Opus 4.6 30.7 / Gemini 3 Pro 37.2 / GPT-5.2 31.4. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).

打开官方来源

scicode 44.4 模型 minimax-m2-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: SciCode · figure: images/84.webp (AA Intelligence Index sub-table)

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

SciCode row: M2.5 44.4 vs M2.1 41.0 / Sonnet 4.5 45.0 / Opus 4.5 50.0 / Opus 4.6 52.0 / Gemini 3 Pro 56.0 / GPT-5.2 52.0. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).

打开官方来源

ifbench 70 模型 minimax-m2-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: IFBench · figure: images/84.webp (AA Intelligence Index sub-table)

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

IFBench row: M2.5 70.0 vs M2.1 70.0 / Sonnet 4.5 57.0 / Opus 4.5 58.0 / Opus 4.6 53.0 / Gemini 3 Pro 70.0 / GPT-5.2 75.0. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).

打开官方来源

aa-lcr 69.5 模型 minimax-m2-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding / Appendix further benchmark results · row: AA-LCR · figure: images/84.webp (AA Intelligence Index sub-table)

{
  "harness": "Claude Code (internal infrastructure, default system prompt overridden)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "4 runs averaged",
  "aggregation": null,
  "judge": null
}

AA-LCR row: M2.5 69.5 vs M2.1 62.0 / Sonnet 4.5 66.0 / Opus 4.5 74.0 / Opus 4.6 71.0 / Gemini 3 Pro 71.0 / GPT-5.2 73.0. Internal testing per public AA Intelligence Index evaluation methods (appendix). Table read (vision, 2026-09-01); competitors in same table: M2.1 / Claude Sonnet 4.5 / Opus 4.5 / Opus 4.6 / Gemini 3 Pro / GPT-5.2 (thinking).

打开官方来源