GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna
OpenAI · 2026-08-26 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GPT-5.6 Sol
发布文将 Sol 定位为 GPT-5.6 的旗舰档(medium/high/xhigh/max/ultra 五档努力级别)。36 项评测横跨智能体、编码、长上下文、科研医疗金融与安全,亮点为 BrowseComp 90.4% 与 Terminal-Bench 2.1 88.8%。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 5 / 输出 30 · 页面另载 Sol 有为期 3 个月、降价逾 20% 的促销记录
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Agents' Last Exam · quote_snippet: GPT-5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: agents-last-exam not yet in data/benchmarks.json. Page describes it as covering long-running professional workflows across 55 fields. Effort setting for the 53.6 figure is not stated; a separate claim says at medium reasoning Sol beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. Competitor deltas are stated but Fable 5's absolute score is not printed on this page. 2026-09-01 audit: page-internal discrepancy - prose prints 53.6 for Sol while the evaluation table prints 52.7 (row value now carries the table cell; prose 53.6 preserved here and in quote_snippet; prose also claims 51.9 at medium reasoning = Fable 40.5 + 11.4).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Artificial Analysis Intelligence Index v4.1 · quote_snippet: GPT-5.6 Sol with max reasoning comes within one point of Fable 5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "tasks completed in 61% less time (relative claim)",
"run_count": null,
"aggregation": "composite index",
"judge": null
}new-benchmark: aa-intelligence-index not yet in data/benchmarks.json. Absolute Sol score not printed on the page - only the relative claim (within one point of Fable 5, 61% less time, roughly half estimated cost). Third-party index quoted by vendor. 2026-09-01 audit: table machine-read - page prints 'Artificial Analysis Intelligence Index v4.1' Sol = 58.9 Index score; previous not_reported status resolved (Fable 5 table cell is 59.9, consistent with the 'within one point' prose).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: coding section · quote_snippet: GPT-5.6 Sol with max reasoning sets a new state of the art at 80, 2.8 points above Fable 5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": "less than half the output tokens of Fable 5 (relative claim)",
"turn_limit": null,
"time_limit": "less than half the time (relative claim)",
"run_count": null,
"aggregation": "composite index",
"judge": null
}new-benchmark: aa-coding-agent-index not yet in data/benchmarks.json. Third-party index quoted by vendor. Same section states Terra performs just above Fable 5 and Luna outperforms Opus 4.8 on this index, each in roughly one-third of the time with about half the output tokens at approximately one-quarter the estimated cost (relative claims, no absolute scores printed).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1 · quote_snippet: It also sets new state-of-the-art results on Terminal-Bench 2.1 and DeepSWE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Qualitative SOTA claim only; no score printed on this release page. A companion page openai.com/index/previewing-gpt-5-6-sol/ is cited by other vendors as the numeric source and should be harvested before publishing any number. 2026-09-01 audit: table machine-read - Terminal-Bench 2.1 Sol = 88.8% (Sol Ultra column 91.9%). The companion-page caveat in earlier notes is resolved for this release page; id kept for continuity though the qualitative-only framing no longer applies.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1 · quote_snippet: sets new state-of-the-art results on Terminal-Bench 2.1 and DeepSWE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: deepswe not yet in data/benchmarks/. Qualitative SOTA claim only; DeepSWE variant (v1.1 or other) not stated on this page. 2026-09-01 audit: table machine-read - DeepSWE row is labelled 'DeepSWE v1.1' on the page (variant now recorded), Sol = 72.7%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp · quote_snippet: GPT-5.6 Sol sets new state-of-the-art results on BrowseComp at 92.2%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks.json. Cross-reference: Kimi K3's page cites this URL as origin for its BrowseComp comparison column. 2026-09-01 audit: page-internal discrepancy - prose attributes 92.2% to Sol while the evaluation table prints Sol 90.4% / Sol Ultra 92.2% (row value now the table Sol cell; 92.2% kept on the Sol Ultra row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: knowledge work section · quote_snippet: OSWorld 2.0 at 62.6%; on OSWorld, it surpasses Opus 4.8 while using 85% fewer output tokens
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "85% fewer output tokens than Opus 4.8 (relative claim)",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark osworld, variant 2.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: cybersecurity section · quote_snippet: it scores 73.5% versus GPT-5.5's 47.9% at a comparable output-token budget
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "comparable output-token budget stated only as a relative condition",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: exploitbench not yet in data/benchmarks/. GPT-5.5 comparison value recorded as a separate comparison_cited row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: cybersecurity section · quote_snippet: it almost doubles GPT-5.5's peak pass rate, from 15.1% to 24.9% under the two-hour cap
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "2h cap",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: exploitgym not yet in data/benchmarks.json. Note: Z.ai's ExploitGym footnote reports task counts under TPS-rescaled budgets, OpenAI reports pass rates - the two vendors use different metrics/aggregations for the same benchmark name; not comparable across these pages.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: cybersecurity section · quote_snippet: with six hours, it reaches 33.7%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "6h",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: exploitgym not yet in data/benchmarks/. Six-hour budget variant of the ExploitGym row above. 2026-09-01 audit: the evaluation-table ExploitGym row prints Sol 33.7% (equals this six-hour figure) alongside Terra 23.2% / Luna 12.4% / GPT-5.5 15.1% - but GPT-5.5's 15.1% equals the two-hour prose figure, so the table's per-column budget is internally inconsistent; Terra/Luna rows are recorded without a budget variant.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: cybersecurity section · quote_snippet: SEC-Bench Pro, which tests proof-of-concept generation ... it scores 71.2%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: sec-bench-pro not yet in data/benchmarks.json. Also used in the multi-agent (ultra) latency charts alongside BrowseComp and Terminal-Bench 2.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Big Finance Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: big-finance-bench not yet in data/benchmarks. Page cells print integer percentages.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: LifeSciBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: lifescibench not yet in data/benchmarks.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。v2 protocol.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 256K-512K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 512K-1M
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: KernelGen 1P
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: kernelgen not yet in data/benchmarks; page label 'KernelGen 1P' (1P variant semantics not explained on page).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: NanoGPT
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: nanogpt not yet in data/benchmarks. Page does not describe the metric; identity vs the public NanoGPT speedrun unconfirmed.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: RSI Index
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: rsi-index not yet in data/benchmarks (recursion/self-improvement index).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: PostTrainBench Lite
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'PostTrainBench Lite' - Lite variant of posttrain-bench.
GPT-5.6 Terra
发布文将 Terra 定位为 GPT-5.6 的均衡档。34 项评测与旗舰同表覆盖智能体、编码、长上下文与专业知识,亮点为 Terminal-Bench 2.1 87.4% 与 SWE-bench Pro 63.4%。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 2.5 / 输出 15
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Big Finance Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: big-finance-bench not yet in data/benchmarks. Page cells print integer percentages.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: LifeSciBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: lifescibench not yet in data/benchmarks.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score').
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OSWorld 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the ExploitBench row (Sol 73.5% / GPT-5.5 47.9% already recorded).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitGym
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Budget not stated in the table row; see the two-hour/six-hour prose rows - table budget is internally inconsistent (see notes there).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SEC-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。v2 protocol.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 256K-512K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 512K-1M
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: KernelGen 1P
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: kernelgen not yet in data/benchmarks; page label 'KernelGen 1P' (1P variant semantics not explained on page).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: NanoGPT
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: nanogpt not yet in data/benchmarks. Page does not describe the metric; identity vs the public NanoGPT speedrun unconfirmed.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: RSI Index
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: rsi-index not yet in data/benchmarks (recursion/self-improvement index).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: PostTrainBench Lite
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'PostTrainBench Lite' - Lite variant of posttrain-bench.
GPT-5.6 Luna
发布文将 Luna 定位为 GPT-5.6 的成本效率档(USD 1/6 每百万 tokens)。34 项评测与旗舰同表覆盖智能体与编码,亮点为 Terminal-Bench 2.1 84.7% 与 SWE-bench Pro 62.7%。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 1 / 输出 6
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Big Finance Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: big-finance-bench not yet in data/benchmarks. Page cells print integer percentages.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: LifeSciBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: lifescibench not yet in data/benchmarks.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score').
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OSWorld 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the ExploitBench row (Sol 73.5% / GPT-5.5 47.9% already recorded).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitGym
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Budget not stated in the table row; see the two-hour/six-hour prose rows - table budget is internally inconsistent (see notes there).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SEC-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。v2 protocol.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 256K-512K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 512K-1M
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: KernelGen 1P
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: kernelgen not yet in data/benchmarks; page label 'KernelGen 1P' (1P variant semantics not explained on page).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: NanoGPT
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: nanogpt not yet in data/benchmarks. Page does not describe the metric; identity vs the public NanoGPT speedrun unconfirmed.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: RSI Index
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: rsi-index not yet in data/benchmarks (recursion/self-improvement index).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: PostTrainBench Lite
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'PostTrainBench Lite' - Lite variant of posttrain-bench.
未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: cybersecurity section · quote_snippet: 73.5% versus GPT-5.5's 47.9%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: exploitbench not yet in data/benchmarks/. Prior-generation competitor score cited by OpenAI in its own release; OpenAI ran or sourced it, not GPT-5.5's own release claim on this page.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: cybersecurity section · quote_snippet: from 15.1% to 24.9% under the two-hour cap
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "2h cap",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: exploitgym not yet in data/benchmarks/. Prior-generation competitor score cited by OpenAI; described as GPT-5.5's peak pass rate under the two-hour cap. 2026-09-01 audit: same 15.1% value also appears in the ExploitGym evaluation-table row for GPT-5.5 (table budget unstated).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: cybersecurity section · quote_snippet: 71.2% versus GPT-5.5's 45.8% at an improved latency
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: sec-bench-pro not yet in data/benchmarks/. Prior-generation competitor score cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Agents' Last Exam
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the Agents' Last Exam table row. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GDPval-AA v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Claude Fable 5 (1,759.6) leads this Elo table; Sol cell is the highest GPT-5.6-tier value. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Big Finance Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: big-finance-bench not yet in data/benchmarks. Page cells print integer percentages. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: Big Finance Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: big-finance-bench not yet in data/benchmarks. Page cells print integer percentages. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: GeneBench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'GeneBench Pro' - recorded as variant Pro of genebench. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: LifeSciBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: lifescibench not yet in data/benchmarks. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: LifeSciBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: lifescibench not yet in data/benchmarks. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: HealthBench Professional
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (6) not reproduced in the row label. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score'). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score'). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score'). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Artificial Analysis Coding Agent Index v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the AA Coding Agent Index row (Sol = 80 already recorded; page unit 'Index score'). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SWE-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。First table rows for SWE-Bench Pro on this release page; Claude Mythos 5 / Fable 5 lead (80.3% / 80%). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Terminal-Bench 2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: DeepSWE v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OSWorld 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OSWorld 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Sol Ultra cell (92.2%) is the number the prose SOTA claim refers to. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: BenchCAD (python tool)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: benchcad not yet in data/benchmarks; python-tool variant row. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the ExploitBench row (Sol 73.5% / GPT-5.5 47.9% already recorded). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the ExploitBench row (Sol 73.5% / GPT-5.5 47.9% already recorded). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ExploitBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the ExploitBench row (Sol 73.5% / GPT-5.5 47.9% already recorded). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - cybersecurity · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: SEC-Bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: MMMU Pro (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - multimodal · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview / Gemini 3.5 Flash · row: gdp.pdf
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 1-3 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page labels the rows 'FrontierMath Tier 1-3 (v2)' / 'Tier 4 (v2)' - v2 protocol. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。v2 protocol. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。v2 protocol. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - academic · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: FrontierMath Tier 4 (v2)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。v2 protocol. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: AutomationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 256K-512K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: OpenAI MRCR v2 8-needle 512K-1M
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: GraphWalks BFS 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Sol Ultra / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 / Claude Mythos 5 / Claude Mythos Preview / Claude Fable 5 / Claude Opus 4.8 / Gemini 3.1 Pro Preview · row: ARC-AGI-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page superscript footnote marker (7) not reproduced in the row label. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: KernelGen 1P
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: kernelgen not yet in data/benchmarks; page label 'KernelGen 1P' (1P variant semantics not explained on page). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: NanoGPT
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: nanogpt not yet in data/benchmarks. Page does not describe the metric; identity vs the public NanoGPT speedrun unconfirmed. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: RSI Index
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: rsi-index not yet in data/benchmarks (recursion/self-improvement index). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - self-improvement · table: text table, columns: GPT-5.6 Sol / GPT-5.6 Terra / GPT-5.6 Luna / GPT-5.5 · row: PostTrainBench Lite
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Page label 'PostTrainBench Lite' - Lite variant of posttrain-bench. Comparison column cited by OpenAI.