← 模型目录

GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna

OpenAI · 2026-08-26 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GPT-5.6 Sol

发布文将 Sol 定位为 GPT-5.6 的旗舰档(medium/high/xhigh/max/ultra 五档努力级别)。36 项评测横跨智能体、编码、长上下文、科研医疗金融与安全,亮点为 BrowseComp 90.4% 与 Terminal-Bench 2.1 88.8%。

输入模态
文本 / 图像
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 5 / 输出 30 · 页面另载 Sol 有为期 3 个月、降价逾 20% 的促销记录

本变体的评测证据

agents-last-exam 52.7% 模型 gpt-5-6-sol · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Agents' Last Exam · quote_snippet: GPT-5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: agents-last-exam not yet in data/benchmarks.json. Page describes it as covering long-running professional workflows across 55 fields. Effort setting for the 53.6 figure is not stated; a separate claim says at medium reasoning Sol beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. Competitor deltas are stated but Fable 5's absolute score is not printed on this page. 2026-09-01 audit: page-internal discrepancy - prose prints 53.6 for Sol while the evaluation table prints 52.7 (row value now carries the table cell; prose 53.6 preserved here and in quote_snippet; prose also claims 51.9 at medium reasoning = Fable 40.5 + 11.4).

打开官方来源

aa-intelligence-index 58.9 模型 gpt-5-6-sol · 版本 未说明 · 指标 composite_intelligence_index · 单位 points 来源等级 A · third_party_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Artificial Analysis Intelligence Index v4.1 · quote_snippet: GPT-5.6 Sol with max reasoning comes within one point of Fable 5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "tasks completed in 61% less time (relative claim)",
  "run_count": null,
  "aggregation": "composite index",
  "judge": null
}

new-benchmark: aa-intelligence-index not yet in data/benchmarks.json. Absolute Sol score not printed on the page - only the relative claim (within one point of Fable 5, 61% less time, roughly half estimated cost). Third-party index quoted by vendor. 2026-09-01 audit: table machine-read - page prints 'Artificial Analysis Intelligence Index v4.1' Sol = 58.9 Index score; previous not_reported status resolved (Fable 5 table cell is 59.9, consistent with the 'within one point' prose).

打开官方来源

aa-coding-agent-index 80 模型 gpt-5-6-sol · 版本 未说明 · 指标 composite_coding_agent_index · 单位 points 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: coding section · quote_snippet: GPT-5.6 Sol with max reasoning sets a new state of the art at 80, 2.8 points above Fable 5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": "less than half the output tokens of Fable 5 (relative claim)",
  "turn_limit": null,
  "time_limit": "less than half the time (relative claim)",
  "run_count": null,
  "aggregation": "composite index",
  "judge": null
}

new-benchmark: aa-coding-agent-index not yet in data/benchmarks.json. Third-party index quoted by vendor. Same section states Terra performs just above Fable 5 and Luna outperforms Opus 4.8 on this index, each in roughly one-third of the time with about half the output tokens at approximately one-quarter the estimated cost (relative claims, no absolute scores printed).

打开官方来源

terminalbench 88.8% 模型 gpt-5-6-sol · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1 · quote_snippet: It also sets new state-of-the-art results on Terminal-Bench 2.1 and DeepSWE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Qualitative SOTA claim only; no score printed on this release page. A companion page openai.com/index/previewing-gpt-5-6-sol/ is cited by other vendors as the numeric source and should be harvested before publishing any number. 2026-09-01 audit: table machine-read - Terminal-Bench 2.1 Sol = 88.8% (Sol Ultra column 91.9%). The companion-page caveat in earlier notes is resolved for this release page; id kept for continuity though the qualitative-only framing no longer applies.

打开官方来源

deepswe 72.7% 模型 gpt-5-6-sol · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1 · quote_snippet: sets new state-of-the-art results on Terminal-Bench 2.1 and DeepSWE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: deepswe not yet in data/benchmarks/. Qualitative SOTA claim only; DeepSWE variant (v1.1 or other) not stated on this page. 2026-09-01 audit: table machine-read - DeepSWE row is labelled 'DeepSWE v1.1' on the page (variant now recorded), Sol = 72.7%.

打开官方来源

browsecomp 90.4% 模型 gpt-5-6-sol · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp · quote_snippet: GPT-5.6 Sol sets new state-of-the-art results on BrowseComp at 92.2%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp not yet in data/benchmarks.json. Cross-reference: Kimi K3's page cites this URL as origin for its BrowseComp comparison column. 2026-09-01 audit: page-internal discrepancy - prose attributes 92.2% to Sol while the evaluation table prints Sol 90.4% / Sol Ultra 92.2% (row value now the table Sol cell; 92.2% kept on the Sol Ultra row).

打开官方来源

osworld 62.6% 模型 gpt-5-6-sol · 版本 2.0 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: knowledge work section · quote_snippet: OSWorld 2.0 at 62.6%; on OSWorld, it surpasses Opus 4.8 while using 85% fewer output tokens

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "85% fewer output tokens than Opus 4.8 (relative claim)",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark osworld, variant 2.0.

打开官方来源

exploitbench 73.5% 模型 gpt-5-6-sol · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: cybersecurity section · quote_snippet: it scores 73.5% versus GPT-5.5's 47.9% at a comparable output-token budget

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "comparable output-token budget stated only as a relative condition",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: exploitbench not yet in data/benchmarks/. GPT-5.5 comparison value recorded as a separate comparison_cited row.

打开官方来源

exploitgym 24.9% 模型 gpt-5-6-sol · 版本 2h cap · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: cybersecurity section · quote_snippet: it almost doubles GPT-5.5's peak pass rate, from 15.1% to 24.9% under the two-hour cap

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "2h cap",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: exploitgym not yet in data/benchmarks.json. Note: Z.ai's ExploitGym footnote reports task counts under TPS-rescaled budgets, OpenAI reports pass rates - the two vendors use different metrics/aggregations for the same benchmark name; not comparable across these pages.

打开官方来源

exploitgym 33.7% 模型 gpt-5-6-sol · 版本 6h · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: cybersecurity section · quote_snippet: with six hours, it reaches 33.7%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "6h",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: exploitgym not yet in data/benchmarks/. Six-hour budget variant of the ExploitGym row above. 2026-09-01 audit: the evaluation-table ExploitGym row prints Sol 33.7% (equals this six-hour figure) alongside Terra 23.2% / Luna 12.4% / GPT-5.5 15.1% - but GPT-5.5's 15.1% equals the two-hour prose figure, so the table's per-column budget is internally inconsistent; Terra/Luna rows are recorded without a budget variant.

打开官方来源

sec-bench-pro 71.2% 模型 gpt-5-6-sol · 版本 未说明 · 指标 poc_generation_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: cybersecurity section · quote_snippet: SEC-Bench Pro, which tests proof-of-concept generation ... it scores 71.2%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: sec-bench-pro not yet in data/benchmarks.json. Also used in the multi-agent (ultra) latency charts alongside BrowseComp and Terminal-Bench 2.1.

打开官方来源

gdpval-aa 1,747.8 Elo 模型 gpt-5-6-sol · 版本 v2 · 指标 elo_rating · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value.

打开官方来源

big-finance-bench 53% 模型 gpt-5-6-sol · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Big Finance Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: big-finance-bench not yet in data/benchmarks. Page cells print integer percentages.

打开官方来源

genebench 28.7% 模型 gpt-5-6-sol · 版本 Pro · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench.

打开官方来源

lifescibench 59.9% 模型 gpt-5-6-sol · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: LifeSciBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: lifescibench not yet in data/benchmarks.

打开官方来源

healthbench 60.5% 模型 gpt-5-6-sol · 版本 Professional · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label.

打开官方来源

swebench-pro 64.6% 模型 gpt-5-6-sol · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%).

打开官方来源

benchcad 70.6% 模型 gpt-5-6-sol · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks.

打开官方来源

benchcad 83.4% 模型 gpt-5-6-sol · 版本 python tool · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row.

打开官方来源

mmmu-pro 83% 模型 gpt-5-6-sol · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mmmu-pro 84.6% 模型 gpt-5-6-sol · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

gdp-pdf 30.7% 模型 gpt-5-6-sol · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

gpqa 94.6% 模型 gpt-5-6-sol · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

frontiermath 89% 模型 gpt-5-6-sol · 版本 Tier 1-3 (v2) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol.

打开官方来源

frontiermath 83% 模型 gpt-5-6-sol · 版本 Tier 4 (v2) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。v2 protocol.

打开官方来源

automationbench 18.1% 模型 gpt-5-6-sol · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

toolathlon 58% 模型 gpt-5-6-sol · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 91.5% 模型 gpt-5-6-sol · 版本 v2 8-needle 256K-512K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 256K-512K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 73.8% 模型 gpt-5-6-sol · 版本 v2 8-needle 512K-1M · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 512K-1M

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

graphwalks 90.7% 模型 gpt-5-6-sol · 版本 BFS 256k f1 · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

graphwalks 77.1% 模型 gpt-5-6-sol · 版本 BFS 1mil f1 · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

arc-agi 7.78% 模型 gpt-5-6-sol · 版本 3 · 指标 score_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label.

打开官方来源

kernelgen 61.1% 模型 gpt-5-6-sol · 版本 1P · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: KernelGen 1P

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: kernelgen not yet in data/benchmarks; page label 'KernelGen 1P' (1P variant semantics not explained on page).

打开官方来源

nanogpt 9.69% 模型 gpt-5-6-sol · 版本 未说明 · 指标 score_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: NanoGPT

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: nanogpt not yet in data/benchmarks. Page does not describe the metric; identity vs the public NanoGPT speedrun unconfirmed.

打开官方来源

rsi-index 57.9% 模型 gpt-5-6-sol · 版本 未说明 · 指标 index · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: RSI Index

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: rsi-index not yet in data/benchmarks (recursion/self-improvement index).

打开官方来源

posttrain-bench 50.3% 模型 gpt-5-6-sol · 版本 Lite · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: PostTrainBench Lite

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'PostTrainBench Lite' - Lite variant of posttrain-bench.

打开官方来源

GPT-5.6 Terra

发布文将 Terra 定位为 GPT-5.6 的均衡档。34 项评测与旗舰同表覆盖智能体、编码、长上下文与专业知识,亮点为 Terminal-Bench 2.1 87.4% 与 SWE-bench Pro 63.4%。

输入模态
文本 / 图像
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 2.5 / 输出 15

本变体的评测证据

agents-last-exam 50.4% 模型 gpt-5-6-terra · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row.

打开官方来源

gdpval-aa 1,593 Elo 模型 gpt-5-6-terra · 版本 v2 · 指标 elo_rating · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value.

打开官方来源

big-finance-bench 51% 模型 gpt-5-6-terra · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Big Finance Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: big-finance-bench not yet in data/benchmarks. Page cells print integer percentages.

打开官方来源

genebench 23.3% 模型 gpt-5-6-terra · 版本 Pro · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench.

打开官方来源

lifescibench 56% 模型 gpt-5-6-terra · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: LifeSciBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: lifescibench not yet in data/benchmarks.

打开官方来源

healthbench 57.7% 模型 gpt-5-6-terra · 版本 Professional · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label.

打开官方来源

aa-coding-agent-index 77.4 模型 gpt-5-6-terra · 版本 未说明 · 指标 composite_coding_agent_index · 单位 points 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score').

打开官方来源

swebench-pro 63.4% 模型 gpt-5-6-terra · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%).

打开官方来源

terminalbench 87.4% 模型 gpt-5-6-terra · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

deepswe 69.6% 模型 gpt-5-6-terra · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

osworld 50.2% 模型 gpt-5-6-terra · 版本 2.0 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OSWorld 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

browsecomp 87.5% 模型 gpt-5-6-terra · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to.

打开官方来源

benchcad 62.3% 模型 gpt-5-6-terra · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks.

打开官方来源

benchcad 78.2% 模型 gpt-5-6-terra · 版本 python tool · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row.

打开官方来源

exploitbench 52.9% 模型 gpt-5-6-terra · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the ExploitBench row (Sol 73.5% / GPT-5.5 47.9% already recorded).

打开官方来源

exploitgym 23.2% 模型 gpt-5-6-terra · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitGym

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Budget not stated in the table row; see the two-hour/six-hour prose rows - table budget is internally inconsistent (see notes there).

打开官方来源

sec-bench-pro 57.7% 模型 gpt-5-6-terra · 版本 未说明 · 指标 poc_generation_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SEC-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mmmu-pro 80.7% 模型 gpt-5-6-terra · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mmmu-pro 82% 模型 gpt-5-6-terra · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

gdp-pdf 24.7% 模型 gpt-5-6-terra · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

gpqa 92.9% 模型 gpt-5-6-terra · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

frontiermath 84.9% 模型 gpt-5-6-terra · 版本 Tier 1-3 (v2) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol.

打开官方来源

frontiermath 68.3% 模型 gpt-5-6-terra · 版本 Tier 4 (v2) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。v2 protocol.

打开官方来源

automationbench 15.2% 模型 gpt-5-6-terra · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

toolathlon 53.1% 模型 gpt-5-6-terra · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 89.6% 模型 gpt-5-6-terra · 版本 v2 8-needle 256K-512K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 256K-512K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 72.5% 模型 gpt-5-6-terra · 版本 v2 8-needle 512K-1M · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 512K-1M

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

graphwalks 76.9% 模型 gpt-5-6-terra · 版本 BFS 256k f1 · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

graphwalks 71.2% 模型 gpt-5-6-terra · 版本 BFS 1mil f1 · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

arc-agi 0.8% 模型 gpt-5-6-terra · 版本 3 · 指标 score_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label.

打开官方来源

kernelgen 49.2% 模型 gpt-5-6-terra · 版本 1P · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: KernelGen 1P

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: kernelgen not yet in data/benchmarks; page label 'KernelGen 1P' (1P variant semantics not explained on page).

打开官方来源

nanogpt 14.5% 模型 gpt-5-6-terra · 版本 未说明 · 指标 score_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: NanoGPT

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: nanogpt not yet in data/benchmarks. Page does not describe the metric; identity vs the public NanoGPT speedrun unconfirmed.

打开官方来源

rsi-index 56.3% 模型 gpt-5-6-terra · 版本 未说明 · 指标 index · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: RSI Index

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: rsi-index not yet in data/benchmarks (recursion/self-improvement index).

打开官方来源

posttrain-bench 51.5% 模型 gpt-5-6-terra · 版本 Lite · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: PostTrainBench Lite

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'PostTrainBench Lite' - Lite variant of posttrain-bench.

打开官方来源

GPT-5.6 Luna

发布文将 Luna 定位为 GPT-5.6 的成本效率档(USD 1/6 每百万 tokens)。34 项评测与旗舰同表覆盖智能体与编码,亮点为 Terminal-Bench 2.1 84.7% 与 SWE-bench Pro 62.7%。

输入模态
文本 / 图像
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 1 / 输出 6

本变体的评测证据

agents-last-exam 50.3% 模型 gpt-5-6-luna · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row.

打开官方来源

gdpval-aa 1,591.8 Elo 模型 gpt-5-6-luna · 版本 v2 · 指标 elo_rating · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value.

打开官方来源

big-finance-bench 36% 模型 gpt-5-6-luna · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Big Finance Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: big-finance-bench not yet in data/benchmarks. Page cells print integer percentages.

打开官方来源

genebench 10.8% 模型 gpt-5-6-luna · 版本 Pro · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench.

打开官方来源

lifescibench 51.2% 模型 gpt-5-6-luna · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: LifeSciBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: lifescibench not yet in data/benchmarks.

打开官方来源

healthbench 55.7% 模型 gpt-5-6-luna · 版本 Professional · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label.

打开官方来源

aa-coding-agent-index 74.6 模型 gpt-5-6-luna · 版本 未说明 · 指标 composite_coding_agent_index · 单位 points 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score').

打开官方来源

swebench-pro 62.7% 模型 gpt-5-6-luna · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%).

打开官方来源

terminalbench 84.7% 模型 gpt-5-6-luna · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

deepswe 67.2% 模型 gpt-5-6-luna · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

osworld 45.6% 模型 gpt-5-6-luna · 版本 2.0 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OSWorld 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

browsecomp 83.3% 模型 gpt-5-6-luna · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to.

打开官方来源

benchcad 63.1% 模型 gpt-5-6-luna · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks.

打开官方来源

benchcad 73.9% 模型 gpt-5-6-luna · 版本 python tool · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row.

打开官方来源

exploitbench 33.2% 模型 gpt-5-6-luna · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the ExploitBench row (Sol 73.5% / GPT-5.5 47.9% already recorded).

打开官方来源

exploitgym 12.4% 模型 gpt-5-6-luna · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitGym

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Budget not stated in the table row; see the two-hour/six-hour prose rows - table budget is internally inconsistent (see notes there).

打开官方来源

sec-bench-pro 48.9% 模型 gpt-5-6-luna · 版本 未说明 · 指标 poc_generation_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SEC-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mmmu-pro 78.4% 模型 gpt-5-6-luna · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mmmu-pro 79.5% 模型 gpt-5-6-luna · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

gdp-pdf 22.7% 模型 gpt-5-6-luna · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

gpqa 92.3% 模型 gpt-5-6-luna · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

frontiermath 78.6% 模型 gpt-5-6-luna · 版本 Tier 1-3 (v2) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol.

打开官方来源

frontiermath 58.5% 模型 gpt-5-6-luna · 版本 Tier 4 (v2) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。v2 protocol.

打开官方来源

automationbench 14.9% 模型 gpt-5-6-luna · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

toolathlon 53.4% 模型 gpt-5-6-luna · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 41.3% 模型 gpt-5-6-luna · 版本 v2 8-needle 256K-512K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 256K-512K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 41.3% 模型 gpt-5-6-luna · 版本 v2 8-needle 512K-1M · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 512K-1M

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

graphwalks 81.3% 模型 gpt-5-6-luna · 版本 BFS 256k f1 · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

graphwalks 51.2% 模型 gpt-5-6-luna · 版本 BFS 1mil f1 · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

arc-agi 0.18% 模型 gpt-5-6-luna · 版本 3 · 指标 score_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label.

打开官方来源

kernelgen 22.4% 模型 gpt-5-6-luna · 版本 1P · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: KernelGen 1P

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: kernelgen not yet in data/benchmarks; page label 'KernelGen 1P' (1P variant semantics not explained on page).

打开官方来源

nanogpt 1.66% 模型 gpt-5-6-luna · 版本 未说明 · 指标 score_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: NanoGPT

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: nanogpt not yet in data/benchmarks. Page does not describe the metric; identity vs the public NanoGPT speedrun unconfirmed.

打开官方来源

rsi-index 41.9% 模型 gpt-5-6-luna · 版本 未说明 · 指标 index · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: RSI Index

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: rsi-index not yet in data/benchmarks (recursion/self-improvement index).

打开官方来源

posttrain-bench 29.6% 模型 gpt-5-6-luna · 版本 Lite · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: PostTrainBench Lite

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'PostTrainBench Lite' - Lite variant of posttrain-bench.

打开官方来源

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

exploitbench 47.9% 模型 gpt-5-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: cybersecurity section · quote_snippet: 73.5% versus GPT-5.5's 47.9%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: exploitbench not yet in data/benchmarks/. Prior-generation competitor score cited by OpenAI in its own release; OpenAI ran or sourced it, not GPT-5.5's own release claim on this page.

打开官方来源

exploitgym 15.1% 模型 gpt-5-5 · 版本 2h cap · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: cybersecurity section · quote_snippet: from 15.1% to 24.9% under the two-hour cap

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "2h cap",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: exploitgym not yet in data/benchmarks/. Prior-generation competitor score cited by OpenAI; described as GPT-5.5's peak pass rate under the two-hour cap. 2026-09-01 audit: same 15.1% value also appears in the ExploitGym evaluation-table row for GPT-5.5 (table budget unstated).

打开官方来源

sec-bench-pro 45.8% 模型 gpt-5-5 · 版本 未说明 · 指标 poc_generation_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: cybersecurity section · quote_snippet: 71.2% versus GPT-5.5's 45.8% at an improved latency

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: sec-bench-pro not yet in data/benchmarks/. Prior-generation competitor score cited by OpenAI.

打开官方来源

agents-last-exam 46.9% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row. Comparison column cited by OpenAI.

打开官方来源

agents-last-exam 40.5% 模型 claude-fable-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row. Comparison column cited by OpenAI.

打开官方来源

agents-last-exam 45.2% 模型 claude-opus-4-8 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row. Comparison column cited by OpenAI.

打开官方来源

agents-last-exam 32.1% 模型 gemini-3-1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row. Comparison column cited by OpenAI.

打开官方来源

gdpval-aa 1,493.7 Elo 模型 gpt-5-5 · 版本 v2 · 指标 elo_rating · 单位 elo 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value. Comparison column cited by OpenAI.

打开官方来源

gdpval-aa 1,759.6 Elo 模型 claude-fable-5 · 版本 v2 · 指标 elo_rating · 单位 elo 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value. Comparison column cited by OpenAI.

打开官方来源

gdpval-aa 1,600.1 Elo 模型 claude-opus-4-8 · 版本 v2 · 指标 elo_rating · 单位 elo 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value. Comparison column cited by OpenAI.

打开官方来源

gdpval-aa 962.3 Elo 模型 gemini-3-1-pro · 版本 v2 · 指标 elo_rating · 单位 elo 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value. Comparison column cited by OpenAI.

打开官方来源

gdpval-aa 1,348.8 Elo 模型 gemini-3-5-flash · 版本 v2 · 指标 elo_rating · 单位 elo 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value. Comparison column cited by OpenAI.

打开官方来源

big-finance-bench 49% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Big Finance Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: big-finance-bench not yet in data/benchmarks. Page cells print integer percentages. Comparison column cited by OpenAI.

打开官方来源

big-finance-bench 44% 模型 claude-opus-4-8 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Big Finance Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: big-finance-bench not yet in data/benchmarks. Page cells print integer percentages. Comparison column cited by OpenAI.

打开官方来源

genebench 12% 模型 gpt-5-5 · 版本 Pro · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench. Comparison column cited by OpenAI.

打开官方来源

genebench 16% 模型 claude-opus-4-8 · 版本 Pro · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench. Comparison column cited by OpenAI.

打开官方来源

genebench 3.1% 模型 gemini-3-1-pro · 版本 Pro · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench. Comparison column cited by OpenAI.

打开官方来源

genebench 8.14% 模型 gemini-3-5-flash · 版本 Pro · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench. Comparison column cited by OpenAI.

打开官方来源

lifescibench 50.4% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: LifeSciBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: lifescibench not yet in data/benchmarks. Comparison column cited by OpenAI.

打开官方来源

lifescibench 53.6% 模型 claude-opus-4-8 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: LifeSciBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: lifescibench not yet in data/benchmarks. Comparison column cited by OpenAI.

打开官方来源

healthbench 49.5% 模型 gpt-5-5 · 版本 Professional · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label. Comparison column cited by OpenAI.

打开官方来源

healthbench 60.9% 模型 claude-fable-5 · 版本 Professional · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label. Comparison column cited by OpenAI.

打开官方来源

healthbench 53% 模型 claude-opus-4-8 · 版本 Professional · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label. Comparison column cited by OpenAI.

打开官方来源

aa-coding-agent-index 76.4 模型 gpt-5-5 · 版本 未说明 · 指标 composite_coding_agent_index · 单位 points 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score'). Comparison column cited by OpenAI.

打开官方来源

aa-coding-agent-index 77.2 模型 claude-fable-5 · 版本 未说明 · 指标 composite_coding_agent_index · 单位 points 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score'). Comparison column cited by OpenAI.

打开官方来源

aa-coding-agent-index 72.5 模型 claude-opus-4-8 · 版本 未说明 · 指标 composite_coding_agent_index · 单位 points 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score'). Comparison column cited by OpenAI.

打开官方来源

aa-coding-agent-index 42.7 模型 gemini-3-1-pro · 版本 未说明 · 指标 composite_coding_agent_index · 单位 points 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score'). Comparison column cited by OpenAI.

打开官方来源

swebench-pro 59.4% 模型 gpt-5-5 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.

打开官方来源

swebench-pro 80.3% 模型 claude-mythos-5 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.

打开官方来源

swebench-pro 77.8% 模型 claude-mythos-preview · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.

打开官方来源

swebench-pro 80% 模型 claude-fable-5 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.

打开官方来源

swebench-pro 69.2% 模型 claude-opus-4-8 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.

打开官方来源

swebench-pro 54.2% 模型 gemini-3-1-pro · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.

打开官方来源

terminalbench 91.9% 模型 gpt-5-6-sol-ultra · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

terminalbench 85.6% 模型 gpt-5-5 · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

terminalbench 88% 模型 claude-mythos-5 · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

terminalbench 83.1% 模型 claude-fable-5 · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

terminalbench 78.9% 模型 claude-opus-4-8 · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

terminalbench 70.7% 模型 gemini-3-1-pro · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

deepswe 67% 模型 gpt-5-5 · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

deepswe 69.7% 模型 claude-fable-5 · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

deepswe 59% 模型 claude-opus-4-8 · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

deepswe 11.8% 模型 gemini-3-1-pro · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

osworld 47.5% 模型 gpt-5-5 · 版本 2.0 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OSWorld 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

osworld 54.8% 模型 claude-opus-4-8 · 版本 2.0 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OSWorld 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

browsecomp 92.2% 模型 gpt-5-6-sol-ultra · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to.

打开官方来源

browsecomp 84.4% 模型 gpt-5-5 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to. Comparison column cited by OpenAI.

打开官方来源

browsecomp 88% 模型 claude-mythos-5 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to. Comparison column cited by OpenAI.

打开官方来源

browsecomp 87.9% 模型 claude-mythos-preview · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to. Comparison column cited by OpenAI.

打开官方来源

browsecomp 84.3% 模型 claude-opus-4-8 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to. Comparison column cited by OpenAI.

打开官方来源

browsecomp 85.9% 模型 gemini-3-1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to. Comparison column cited by OpenAI.

打开官方来源

benchcad 44.4% 模型 gpt-5-5 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks. Comparison column cited by OpenAI.

打开官方来源

benchcad 38.4% 模型 claude-mythos-5 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks. Comparison column cited by OpenAI.

打开官方来源

benchcad 35.5% 模型 claude-mythos-preview · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks. Comparison column cited by OpenAI.

打开官方来源

benchcad 27.3% 模型 claude-opus-4-8 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks. Comparison column cited by OpenAI.

打开官方来源

benchcad 55.8% 模型 gpt-5-5 · 版本 python tool · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row. Comparison column cited by OpenAI.

打开官方来源

benchcad 65% 模型 claude-mythos-5 · 版本 python tool · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row. Comparison column cited by OpenAI.

打开官方来源

benchcad 61% 模型 claude-mythos-preview · 版本 python tool · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row. Comparison column cited by OpenAI.

打开官方来源

benchcad 51.8% 模型 claude-opus-4-8 · 版本 python tool · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row. Comparison column cited by OpenAI.

打开官方来源

exploitbench 78% 模型 claude-mythos-5 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the ExploitBench row (Sol 73.5% / GPT-5.5 47.9% already recorded). Comparison column cited by OpenAI.

打开官方来源

exploitbench 74.2% 模型 claude-mythos-preview · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the ExploitBench row (Sol 73.5% / GPT-5.5 47.9% already recorded). Comparison column cited by OpenAI.

打开官方来源

exploitbench 40% 模型 claude-opus-4-8 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the ExploitBench row (Sol 73.5% / GPT-5.5 47.9% already recorded). Comparison column cited by OpenAI.

打开官方来源

sec-bench-pro 74.3% 模型 gpt-5-6-sol-ultra · 版本 未说明 · 指标 poc_generation_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SEC-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mmmu-pro 81.2% 模型 gpt-5-5 · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mmmu-pro 80.5% 模型 gemini-3-1-pro · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mmmu-pro 83.2% 模型 gpt-5-5 · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gdp-pdf 26% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gdp-pdf 29.8% 模型 claude-fable-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gdp-pdf 22.5% 模型 claude-opus-4-8 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gdp-pdf 16.7% 模型 gemini-3-1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gpqa 93.6% 模型 gpt-5-5 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gpqa 94.1% 模型 claude-mythos-5 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gpqa 94.6% 模型 claude-mythos-preview · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gpqa 92.6% 模型 claude-fable-5 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gpqa 92% 模型 claude-opus-4-8 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gpqa 94.3% 模型 gemini-3-1-pro · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

frontiermath 85.3% 模型 gpt-5-5 · 版本 Tier 1-3 (v2) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol. Comparison column cited by OpenAI.

打开官方来源

frontiermath 87% 模型 claude-fable-5 · 版本 Tier 1-3 (v2) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol. Comparison column cited by OpenAI.

打开官方来源

frontiermath 80% 模型 claude-opus-4-8 · 版本 Tier 1-3 (v2) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol. Comparison column cited by OpenAI.

打开官方来源

frontiermath 59.6% 模型 gemini-3-1-pro · 版本 Tier 1-3 (v2) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol. Comparison column cited by OpenAI.

打开官方来源

frontiermath 72.5% 模型 gpt-5-5 · 版本 Tier 4 (v2) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。v2 protocol. Comparison column cited by OpenAI.

打开官方来源

frontiermath 87.8% 模型 claude-fable-5 · 版本 Tier 4 (v2) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。v2 protocol. Comparison column cited by OpenAI.

打开官方来源

frontiermath 56.1% 模型 claude-opus-4-8 · 版本 Tier 4 (v2) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。v2 protocol. Comparison column cited by OpenAI.

打开官方来源

automationbench 12.9% 模型 gpt-5-5 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

automationbench 17.4% 模型 claude-fable-5 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

automationbench 15.5% 模型 claude-opus-4-8 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

automationbench 14.5% 模型 gemini-3-5-flash · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

toolathlon 55.6% 模型 gpt-5-5 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

toolathlon 61.7% 模型 claude-mythos-5 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

toolathlon 61.1% 模型 claude-mythos-preview · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

toolathlon 61.7% 模型 claude-fable-5 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

toolathlon 59.9% 模型 claude-opus-4-8 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

toolathlon 48.8% 模型 gemini-3-1-pro · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mrcr 81.5% 模型 gpt-5-5 · 版本 v2 8-needle 256K-512K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 256K-512K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mrcr 74% 模型 gpt-5-5 · 版本 v2 8-needle 512K-1M · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 512K-1M

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

graphwalks 73.7% 模型 gpt-5-5 · 版本 BFS 256k f1 · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

graphwalks 91.1% 模型 claude-mythos-5 · 版本 BFS 256k f1 · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

graphwalks 85.7% 模型 claude-mythos-preview · 版本 BFS 256k f1 · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

graphwalks 85.9% 模型 claude-opus-4-8 · 版本 BFS 256k f1 · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

graphwalks 45.4% 模型 gpt-5-5 · 版本 BFS 1mil f1 · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

graphwalks 79.4% 模型 claude-mythos-5 · 版本 BFS 1mil f1 · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

graphwalks 74.3% 模型 claude-mythos-preview · 版本 BFS 1mil f1 · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

graphwalks 68.1% 模型 claude-opus-4-8 · 版本 BFS 1mil f1 · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

arc-agi 0.43% 模型 gpt-5-5 · 版本 3 · 指标 score_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label. Comparison column cited by OpenAI.

打开官方来源

arc-agi 1.5% 模型 claude-opus-4-8 · 版本 3 · 指标 score_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label. Comparison column cited by OpenAI.

打开官方来源

arc-agi 0.42% 模型 gemini-3-1-pro · 版本 3 · 指标 score_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label. Comparison column cited by OpenAI.

打开官方来源

kernelgen 29.3% 模型 gpt-5-5 · 版本 1P · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: KernelGen 1P

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: kernelgen not yet in data/benchmarks; page label 'KernelGen 1P' (1P variant semantics not explained on page). Comparison column cited by OpenAI.

打开官方来源

nanogpt 2.65% 模型 gpt-5-5 · 版本 未说明 · 指标 score_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: NanoGPT

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: nanogpt not yet in data/benchmarks. Page does not describe the metric; identity vs the public NanoGPT speedrun unconfirmed. Comparison column cited by OpenAI.

打开官方来源

rsi-index 41.7% 模型 gpt-5-5 · 版本 未说明 · 指标 index · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: RSI Index

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: rsi-index not yet in data/benchmarks (recursion/self-improvement index). Comparison column cited by OpenAI.

打开官方来源

posttrain-bench 38.8% 模型 gpt-5-5 · 版本 Lite · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: PostTrainBench Lite

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Page label 'PostTrainBench Lite' - Lite variant of posttrain-bench. Comparison column cited by OpenAI.

打开官方来源