GPT-5.5 / GPT-5.5 Pro
OpenAI · 2026-04-23 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GPT-5.5
OpenAI 发布 GPT-5.5,提供 low/medium/high/xhigh 推理档(另有非推理模式)。评测横跨智能体编码、知识工作、长上下文与科学/网络安全,亮点如 Terminal-Bench 2.0 82.7%、ARC-AGI-2 85.0%。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 5 / 输出 30 · 发布时 API 状态为 coming soon;Codex fast mode 为 1.5 倍速、2.5 倍价格
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: SWE-Bench Pro (Public)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}swebench-pro id already introduced by a prior batch; reused. Prose corroborates 58.6%. Page footnote: the lab has found evidence of memorization in this eval - treat scores with caution.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Terminal-Bench 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: "GPT-5.5 achieves a top accuracy of 82.7% on Terminal-Bench 2.0".
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Expert-SWE (Internal)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: expert-swe not yet in data/benchmarks/. Internal frontier eval for long-horizon coding tasks (human median completion ~20 hours). Prose claims GPT-5.5 beats GPT-5.4 without printing a number; the table prints 73.1%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}gdpval id already introduced by a prior batch; reused. Prose corroborates 84.9% (economically valuable knowledge work across 44 professions). Metric is win-or-tie rate, not plain accuracy. A separate vega chart decomposes GDPval into wins 68.2% / ties 16.7% for GPT-5.5 (GPT-5.4 wins 70.8%, GPT-5.4 Pro 69.2%, Opus 4.7 62.6%, Gemini 3.1 Pro 46.8%).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FinanceAgent v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}finance-agent id already introduced by a prior batch; reused. Prose corroborates 60.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OfficeQA Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}officeqa-pro id already introduced by a prior batch; reused. Prose corroborates 54.1%. officeqa-pro id was introduced by an earlier batch (per official/README reuse list) and is not yet migrated into data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OSWorld-Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "xhigh (headline); vega chart by effort: none 52.91 / low 70.96 / medium 76.41 / high 77.53 / xhigh 78.68",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose corroborates 78.7%. Companion vega chart plots accuracy and tool-call count vs effort for GPT-5.5 and GPT-5.4 (tool calls 6.31-12.27 for GPT-5.5).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MMMU Pro (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MMMU Pro (with tools)
{
"harness": null,
"tools": "allowed (page does not enumerate)",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MCP Atlas
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}mcp-atlas id already introduced by a prior batch; reused. Footnote: Scale AI results as of the April 2026 update.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}toolathlon id already introduced by a prior batch; reused.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Tau2-bench Telecom (original prompt)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "user simulator: GPT-4.1 (per page note)"
}new-benchmark: tau2-bench (tau2/t2-bench) not yet in data/benchmarks/ as its own id; distinct from the existing tau-bench entry (retail/airline). Footnote: GPT-5.5 and GPT-5.4 run with original prompts, i.e. no prompt tuning; the page deliberately ignores other labs prompt-tuned results for comparability. Prose corroborates 98.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GeneBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: genebench not yet in data/benchmarks/. New eval for multi-stage scientific data analysis in genetics and quantitative biology; tasks correspond to days of expert work. Prose claims a leap over GPT-5.4 without a number; table prints 25.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BixBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: bixbench not yet in data/benchmarks/. Real bioinformatics / data-analysis benchmark. Separate vega chart gives GPT-5.5 80.5 with CI [79.7, 81.4] vs GPT-5.4 74.0. Prose: ranks at the top among all models with published scores.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)
{
"harness": null,
"tools": "allowed (page does not enumerate)",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Cybersecurity · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: CyberGym
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}cybergym id already introduced by a prior batch; reused. Prose: defensive cyber capability rated high under the preparedness framework; GPT-5.5 ships with stricter risk classifiers and a Trusted Access for Cyber program.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: graphwalks not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: graphwalks not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: graphwalks not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: graphwalks not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 512K-1M
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mrcr not yet in data/benchmarks/. Full curve for GPT-5.5: 4K-8K 98.1 / 8K-16K 93.0 / 16K-32K 96.5 / 32K-64K 90.0 / 64K-128K 83.1 / 128K-256K 87.5 / 256K-512K 81.5 / 512K-1M 74.0. Cross-vendor note: Anthropic reports the same MRCR v2 8-needle family at the 1M tier on the Claude Opus 4.6 page (Opus 4.6 76%, Sonnet 4.5 18.5%).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-1 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-2 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model capabilities section (vega chart) + footnote · figure: vega chart of Artificial Analysis Intelligence Index; values machine-read from page RSC payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "xhigh",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}aa-intelligence-index id already used by gpt-5-6.json; reused. Footnote: index is a third-party weighted average over 10 evals (AA-LCR, AA-Omniscience, CritPt, GDPval-AA, GPQA Diamond, HLE, IFBench, SciCode, Terminal-Bench Hard, tau2-Bench Telecom). Same chart: GPT-5.5 high 58.9 / medium 56.7 / low 50.8 / non-reasoning 40.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 4K-8K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 8K-16K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 16K-32K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 32K-64K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 64K-128K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 128K-256K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 256K-512K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。
GPT-5.5 Pro
GPT-5.5 Pro 为同公告中的最高精度档(highest accuracy tier)。在 BrowseComp 90.1%、HLE with tools 57.2% 等工具型评测上高于标准档。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 30 / 输出 180 · 发布时 API 状态为 coming soon
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GeneBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: genebench not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: SWE-Bench Pro (Public)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation score cited by OpenAI. Memorization footnote applies to this eval.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: SWE-Bench Pro (Public)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor score cited by OpenAI; Opus 4.7 outperforms GPT-5.5 on this row. Memorization footnote applies to this eval.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: SWE-Bench Pro (Public)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor score cited by OpenAI. Memorization footnote applies to this eval.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Terminal-Bench 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Terminal-Bench 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Terminal-Bench 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Expert-SWE (Internal)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: expert-swe not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FinanceAgent v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FinanceAgent v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FinanceAgent v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor score cited by OpenAI; Opus 4.7 outperforms GPT-5.5 on this row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FinanceAgent v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OfficeQA Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}officeqa-pro id was introduced by an earlier batch (per official/README reuse list) and is not yet migrated into data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OfficeQA Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}officeqa-pro id was introduced by an earlier batch (per official/README reuse list) and is not yet migrated into data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OfficeQA Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}officeqa-pro id was introduced by an earlier batch (per official/README reuse list) and is not yet migrated into data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OSWorld-Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vega effort curve for GPT-5.4: none 39.77 / low 69.36 / medium 71.77 / high 73.71 / xhigh 75.03.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OSWorld-Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MMMU Pro (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MMMU Pro (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MCP Atlas
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MCP Atlas
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor score cited by OpenAI; Opus 4.7 outperforms GPT-5.5 on this row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MCP Atlas
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Tau2-bench Telecom (original prompt)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tau2-bench; see vendor row note for protocol.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GeneBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: genebench not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GeneBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: genebench not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BixBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: bixbench not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor score cited by OpenAI; Opus 4.7 outperforms GPT-5.5 on this row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Cybersecurity · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: CyberGym
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Cybersecurity · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: CyberGym
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: graphwalks not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: graphwalks; competitor outperforms GPT-5.5 on this row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: graphwalks not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Table cell explicitly labels this value as Opus 4.6 (not 4.7). new-benchmark: graphwalks not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: graphwalks not yet in data/benchmarks/ (prior batches have not migrated it).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 256k f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: graphwalks; competitor outperforms GPT-5.5 on this row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: graphwalks not yet in data/benchmarks/ (prior batches have not migrated it).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 1mil f1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Table cell explicitly labels this value as Opus 4.6 (not 4.7); competitor outperforms GPT-5.5 on this row. new-benchmark: graphwalks not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 512K-1M
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mrcr not yet in data/benchmarks/. GPT-5.4 full curve: 4K-8K 97.3 / 8K-16K 91.4 / 16K-32K 97.2 / 32K-64K 90.5 / 64K-128K 86.0 / 128K-256K 79.3 / 256K-512K 57.5 / 512K-1M 36.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-1 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-1 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-1 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-1 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor score cited by OpenAI; Gemini 3.1 Pro outperforms GPT-5.5 on this row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-2 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-2 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-2 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-2 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model capabilities section (vega chart) · figure: vega chart; values machine-read from page RSC payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Same chart also lists Gemini 3.1 Pro Preview 57.2, GPT-5.4 xhigh 56.8, Claude Opus 4.6 (max) 53.0, Opus 4.7 non-reasoning 51.8, Opus 4.6 non-reasoning 46.5, GPT-5.4 non-reasoning 35.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 4K-8K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 8K-16K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 16K-32K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 32K-64K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 64K-128K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 128K-256K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 256K-512K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 128K-256K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 512K-1M
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MMMU Pro (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.