← 模型目录

GPT-5.5 / GPT-5.5 Pro

OpenAI · 2026-04-23 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GPT-5.5

OpenAI 发布 GPT-5.5,提供 low/medium/high/xhigh 推理档(另有非推理模式)。评测横跨智能体编码、知识工作、长上下文与科学/网络安全,亮点如 Terminal-Bench 2.0 82.7%、ARC-AGI-2 85.0%。

输入模态
文本 / 图像
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 5 / 输出 30 · 发布时 API 状态为 coming soon;Codex fast mode 为 1.5 倍速、2.5 倍价格

本变体的评测证据

swebench-pro 58.6% 模型 gpt-5-5 · 版本 Public · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: SWE-Bench Pro (Public)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

swebench-pro id already introduced by a prior batch; reused. Prose corroborates 58.6%. Page footnote: the lab has found evidence of memorization in this eval - treat scores with caution.

打开官方来源

terminalbench 82.7% 模型 gpt-5-5 · 版本 2.0 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Terminal-Bench 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: "GPT-5.5 achieves a top accuracy of 82.7% on Terminal-Bench 2.0".

打开官方来源

expert-swe 73.1% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Expert-SWE (Internal)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: expert-swe not yet in data/benchmarks/. Internal frontier eval for long-horizon coding tasks (human median completion ~20 hours). Prose claims GPT-5.5 beats GPT-5.4 without printing a number; the table prints 73.1%.

打开官方来源

gdpval 84.9% 模型 gpt-5-5 · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

gdpval id already introduced by a prior batch; reused. Prose corroborates 84.9% (economically valuable knowledge work across 44 professions). Metric is win-or-tie rate, not plain accuracy. A separate vega chart decomposes GDPval into wins 68.2% / ties 16.7% for GPT-5.5 (GPT-5.4 wins 70.8%, GPT-5.4 Pro 69.2%, Opus 4.7 62.6%, Gemini 3.1 Pro 46.8%).

打开官方来源

finance-agent 60.0% 模型 gpt-5-5 · 版本 v1.1 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FinanceAgent v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

finance-agent id already introduced by a prior batch; reused. Prose corroborates 60.0%.

打开官方来源

officeqa-pro 54.1% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OfficeQA Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

officeqa-pro id already introduced by a prior batch; reused. Prose corroborates 54.1%. officeqa-pro id was introduced by an earlier batch (per official/README reuse list) and is not yet migrated into data/benchmarks/.

打开官方来源

osworld 78.7% 模型 gpt-5-5 · 版本 Verified · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OSWorld-Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "xhigh (headline); vega chart by effort: none 52.91 / low 70.96 / medium 76.41 / high 77.53 / xhigh 78.68",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose corroborates 78.7%. Companion vega chart plots accuracy and tool-call count vs effort for GPT-5.5 and GPT-5.4 (tool calls 6.31-12.27 for GPT-5.5).

打开官方来源

mmmu-pro 81.2% 模型 gpt-5-5 · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MMMU Pro (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).

打开官方来源

mmmu-pro 83.2% 模型 gpt-5-5 · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MMMU Pro (with tools)

{
  "harness": null,
  "tools": "allowed (page does not enumerate)",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).

打开官方来源

browsecomp 84.4% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

mcp-atlas 75.3% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MCP Atlas

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

mcp-atlas id already introduced by a prior batch; reused. Footnote: Scale AI results as of the April 2026 update.

打开官方来源

toolathlon 55.6% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

toolathlon id already introduced by a prior batch; reused.

打开官方来源

tau2-bench 98.0% 模型 gpt-5-5 · 版本 Telecom, original prompts · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Tau2-bench Telecom (original prompt)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "user simulator: GPT-4.1 (per page note)"
}

new-benchmark: tau2-bench (tau2/t2-bench) not yet in data/benchmarks/ as its own id; distinct from the existing tau-bench entry (retail/airline). Footnote: GPT-5.5 and GPT-5.4 run with original prompts, i.e. no prompt tuning; the page deliberately ignores other labs prompt-tuned results for comparability. Prose corroborates 98.0%.

打开官方来源

genebench 25.0% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GeneBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: genebench not yet in data/benchmarks/. New eval for multi-stage scientific data analysis in genetics and quantitative biology; tasks correspond to days of expert work. Prose claims a leap over GPT-5.4 without a number; table prints 25.0%.

打开官方来源

frontiermath 51.7% 模型 gpt-5-5 · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 35.4% 模型 gpt-5-5 · 版本 Tier 4 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

bixbench 80.5% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BixBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: bixbench not yet in data/benchmarks/. Real bioinformatics / data-analysis benchmark. Separate vega chart gives GPT-5.5 80.5 with CI [79.7, 81.4] vs GPT-5.4 74.0. Prose: ranks at the top among all models with published scores.

打开官方来源

gpqa 93.6% 模型 gpt-5-5 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 41.4% 模型 gpt-5-5 · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 52.2% 模型 gpt-5-5 · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)

{
  "harness": null,
  "tools": "allowed (page does not enumerate)",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

cybergym 81.8% 模型 gpt-5-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Cybersecurity · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: CyberGym

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

cybergym id already introduced by a prior batch; reused. Prose: defensive cyber capability rated high under the preparedness framework; GPT-5.5 ships with stricter risk classifiers and a Trusted Access for Cyber program.

打开官方来源

graphwalks 73.7% 模型 gpt-5-5 · 版本 BFS 256k · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: graphwalks not yet in data/benchmarks/.

打开官方来源

graphwalks 45.4% 模型 gpt-5-5 · 版本 BFS 1mil · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: graphwalks not yet in data/benchmarks/.

打开官方来源

graphwalks 90.1% 模型 gpt-5-5 · 版本 parents 256k · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: graphwalks not yet in data/benchmarks/.

打开官方来源

graphwalks 58.5% 模型 gpt-5-5 · 版本 parents 1mil · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: graphwalks not yet in data/benchmarks/.

打开官方来源

mrcr 74.0% 模型 gpt-5-5 · 版本 v2 8-needle 512K-1M · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 512K-1M

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mrcr not yet in data/benchmarks/. Full curve for GPT-5.5: 4K-8K 98.1 / 8K-16K 93.0 / 16K-32K 96.5 / 32K-64K 90.0 / 64K-128K 83.1 / 128K-256K 87.5 / 256K-512K 81.5 / 512K-1M 74.0. Cross-vendor note: Anthropic reports the same MRCR v2 8-needle family at the 1M tier on the Claude Opus 4.6 page (Opus 4.6 76%, Sonnet 4.5 18.5%).

打开官方来源

arc-agi 95.0% 模型 gpt-5-5 · 版本 1 (Verified) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-1 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

arc-agi 85.0% 模型 gpt-5-5 · 版本 2 (Verified) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-2 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

aa-intelligence-index 60.2 模型 gpt-5-5 · 版本 未说明 · 指标 composite_intelligence_index · 单位 points 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model capabilities section (vega chart) + footnote · figure: vega chart of Artificial Analysis Intelligence Index; values machine-read from page RSC payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "xhigh",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

aa-intelligence-index id already used by gpt-5-6.json; reused. Footnote: index is a third-party weighted average over 10 evals (AA-LCR, AA-Omniscience, CritPt, GDPval-AA, GPQA Diamond, HLE, IFBench, SciCode, Terminal-Bench Hard, tau2-Bench Telecom). Same chart: GPT-5.5 high 58.9 / medium 56.7 / low 50.8 / non-reasoning 40.9.

打开官方来源

mrcr 98.1% 模型 gpt-5-5 · 版本 v2 8-needle 4K-8K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 4K-8K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。

打开官方来源

mrcr 93.0% 模型 gpt-5-5 · 版本 v2 8-needle 8K-16K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 8K-16K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。

打开官方来源

mrcr 96.5% 模型 gpt-5-5 · 版本 v2 8-needle 16K-32K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 16K-32K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。

打开官方来源

mrcr 90.0% 模型 gpt-5-5 · 版本 v2 8-needle 32K-64K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 32K-64K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。

打开官方来源

mrcr 83.1% 模型 gpt-5-5 · 版本 v2 8-needle 64K-128K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 64K-128K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。

打开官方来源

mrcr 87.5% 模型 gpt-5-5 · 版本 v2 8-needle 128K-256K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 128K-256K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。

打开官方来源

mrcr 81.5% 模型 gpt-5-5 · 版本 v2 8-needle 256K-512K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 256K-512K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。

打开官方来源

GPT-5.5 Pro

GPT-5.5 Pro 为同公告中的最高精度档(highest accuracy tier)。在 BrowseComp 90.1%、HLE with tools 57.2% 等工具型评测上高于标准档。

输入模态
文本 / 图像
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 30 / 输出 180 · 发布时 API 状态为 coming soon

本变体的评测证据

gdpval 82.3% 模型 gpt-5-5-pro · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

browsecomp 90.1% 模型 gpt-5-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

genebench 33.2% 模型 gpt-5-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GeneBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: genebench not yet in data/benchmarks/.

打开官方来源

frontiermath 52.4% 模型 gpt-5-5-pro · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 39.6% 模型 gpt-5-5-pro · 版本 Tier 4 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 43.1% 模型 gpt-5-5-pro · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 57.2% 模型 gpt-5-5-pro · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

swebench-pro 57.7% 模型 gpt-5-4 · 版本 Public · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: SWE-Bench Pro (Public)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation score cited by OpenAI. Memorization footnote applies to this eval.

打开官方来源

swebench-pro 64.3% 模型 claude-opus-4-7 · 版本 Public · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: SWE-Bench Pro (Public)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor score cited by OpenAI; Opus 4.7 outperforms GPT-5.5 on this row. Memorization footnote applies to this eval.

打开官方来源

swebench-pro 54.2% 模型 gemini-3-1-pro · 版本 Public · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: SWE-Bench Pro (Public)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor score cited by OpenAI. Memorization footnote applies to this eval.

打开官方来源

terminalbench 75.1% 模型 gpt-5-4 · 版本 2.0 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Terminal-Bench 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

terminalbench 69.4% 模型 claude-opus-4-7 · 版本 2.0 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Terminal-Bench 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

terminalbench 68.5% 模型 gemini-3-1-pro · 版本 2.0 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Terminal-Bench 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

expert-swe 68.5% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Coding · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Expert-SWE (Internal)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: expert-swe not yet in data/benchmarks/.

打开官方来源

gdpval 83.0% 模型 gpt-5-4 · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

gdpval 82.0% 模型 gpt-5-4-pro · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

gdpval 80.3% 模型 claude-opus-4-7 · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

gdpval 67.3% 模型 gemini-3-1-pro · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GDPval (win or tie)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

finance-agent 56.0% 模型 gpt-5-4 · 版本 v1.1 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FinanceAgent v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

finance-agent 61.5% 模型 gpt-5-4-pro · 版本 v1.1 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FinanceAgent v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

finance-agent 64.4% 模型 claude-opus-4-7 · 版本 v1.1 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FinanceAgent v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor score cited by OpenAI; Opus 4.7 outperforms GPT-5.5 on this row.

打开官方来源

finance-agent 59.7% 模型 gemini-3-1-pro · 版本 v1.1 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FinanceAgent v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

officeqa-pro 53.2% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OfficeQA Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

officeqa-pro id was introduced by an earlier batch (per official/README reuse list) and is not yet migrated into data/benchmarks/.

打开官方来源

officeqa-pro 43.6% 模型 claude-opus-4-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OfficeQA Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

officeqa-pro id was introduced by an earlier batch (per official/README reuse list) and is not yet migrated into data/benchmarks/.

打开官方来源

officeqa-pro 18.1% 模型 gemini-3-1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Professional · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OfficeQA Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

officeqa-pro id was introduced by an earlier batch (per official/README reuse list) and is not yet migrated into data/benchmarks/.

打开官方来源

osworld 75.0% 模型 gpt-5-4 · 版本 Verified · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OSWorld-Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vega effort curve for GPT-5.4: none 39.77 / low 69.36 / medium 71.77 / high 73.71 / xhigh 75.03.

打开官方来源

osworld 78.0% 模型 claude-opus-4-7 · 版本 Verified · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OSWorld-Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

mmmu-pro 80.5% 模型 gemini-3-1-pro · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MMMU Pro (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).

打开官方来源

mmmu-pro 82.1% 模型 gpt-5-4 · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MMMU Pro (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).

打开官方来源

browsecomp 82.7% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

browsecomp 89.3% 模型 gpt-5-4-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

browsecomp 79.3% 模型 claude-opus-4-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

browsecomp 85.9% 模型 gemini-3-1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

mcp-atlas 70.6% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MCP Atlas

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

mcp-atlas 79.1% 模型 claude-opus-4-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MCP Atlas

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor score cited by OpenAI; Opus 4.7 outperforms GPT-5.5 on this row.

打开官方来源

mcp-atlas 78.2% 模型 gemini-3-1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MCP Atlas

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

toolathlon 54.6% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

toolathlon 48.8% 模型 gemini-3-1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

tau2-bench 92.8% 模型 gpt-5-4 · 版本 Telecom, original prompts · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Tool use · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Tau2-bench Telecom (original prompt)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: tau2-bench; see vendor row note for protocol.

打开官方来源

genebench 19.0% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GeneBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: genebench not yet in data/benchmarks/.

打开官方来源

genebench 25.6% 模型 gpt-5-4-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GeneBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: genebench not yet in data/benchmarks/.

打开官方来源

frontiermath 47.6% 模型 gpt-5-4 · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 50.0% 模型 gpt-5-4-pro · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 43.8% 模型 claude-opus-4-7 · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 36.9% 模型 gemini-3-1-pro · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 1-3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 27.1% 模型 gpt-5-4 · 版本 Tier 4 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 38.0% 模型 gpt-5-4-pro · 版本 Tier 4 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 22.9% 模型 claude-opus-4-7 · 版本 Tier 4 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 16.7% 模型 gemini-3-1-pro · 版本 Tier 4 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: FrontierMath Tier 4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

bixbench 74.0% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: BixBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: bixbench not yet in data/benchmarks/.

打开官方来源

gpqa 92.8% 模型 gpt-5-4 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

gpqa 94.4% 模型 gpt-5-4-pro · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

gpqa 94.2% 模型 claude-opus-4-7 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

gpqa 94.3% 模型 gemini-3-1-pro · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 39.8% 模型 gpt-5-4 · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 42.7% 模型 gpt-5-4-pro · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 46.9% 模型 claude-opus-4-7 · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor score cited by OpenAI; Opus 4.7 outperforms GPT-5.5 on this row.

打开官方来源

hlehle 44.4% 模型 gemini-3-1-pro · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 52.1% 模型 gpt-5-4 · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 58.7% 模型 gpt-5-4-pro · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 54.7% 模型 claude-opus-4-7 · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 51.4% 模型 gemini-3-1-pro · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Academic · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Humanity's Last Exam (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

cybergym 79.0% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Cybersecurity · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: CyberGym

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

cybergym 73.1% 模型 claude-opus-4-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Cybersecurity · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: CyberGym

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

graphwalks 62.5% 模型 gpt-5-4 · 版本 BFS 256k · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: graphwalks not yet in data/benchmarks/.

打开官方来源

graphwalks 76.9% 模型 claude-opus-4-7 · 版本 BFS 256k · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: graphwalks; competitor outperforms GPT-5.5 on this row.

打开官方来源

graphwalks 9.4% 模型 gpt-5-4 · 版本 BFS 1mil · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: graphwalks not yet in data/benchmarks/.

打开官方来源

graphwalks 41.2% 模型 claude-opus-4-6 · 版本 BFS 1mil · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks BFS 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Table cell explicitly labels this value as Opus 4.6 (not 4.7). new-benchmark: graphwalks not yet in data/benchmarks/.

打开官方来源

graphwalks 82.8% 模型 gpt-5-4 · 版本 parents 256k · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: graphwalks not yet in data/benchmarks/ (prior batches have not migrated it).

打开官方来源

graphwalks 93.6% 模型 claude-opus-4-7 · 版本 parents 256k · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 256k f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: graphwalks; competitor outperforms GPT-5.5 on this row.

打开官方来源

graphwalks 44.4% 模型 gpt-5-4 · 版本 parents 1mil · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: graphwalks not yet in data/benchmarks/ (prior batches have not migrated it).

打开官方来源

graphwalks 72.0% 模型 claude-opus-4-6 · 版本 parents 1mil · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: Graphwalks parents 1mil f1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Table cell explicitly labels this value as Opus 4.6 (not 4.7); competitor outperforms GPT-5.5 on this row. new-benchmark: graphwalks not yet in data/benchmarks/.

打开官方来源

mrcr 36.6% 模型 gpt-5-4 · 版本 v2 8-needle 512K-1M · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 512K-1M

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mrcr not yet in data/benchmarks/. GPT-5.4 full curve: 4K-8K 97.3 / 8K-16K 91.4 / 16K-32K 97.2 / 32K-64K 90.5 / 64K-128K 86.0 / 128K-256K 79.3 / 256K-512K 57.5 / 512K-1M 36.6.

打开官方来源

arc-agi 93.7% 模型 gpt-5-4 · 版本 1 (Verified) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-1 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

arc-agi 94.5% 模型 gpt-5-4-pro · 版本 1 (Verified) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-1 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

arc-agi 93.5% 模型 claude-opus-4-7 · 版本 1 (Verified) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-1 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

arc-agi 98.0% 模型 gemini-3-1-pro · 版本 1 (Verified) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-1 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor score cited by OpenAI; Gemini 3.1 Pro outperforms GPT-5.5 on this row.

打开官方来源

arc-agi 73.3% 模型 gpt-5-4 · 版本 2 (Verified) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-2 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

arc-agi 83.3% 模型 gpt-5-4-pro · 版本 2 (Verified) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-2 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

arc-agi 75.8% 模型 claude-opus-4-7 · 版本 2 (Verified) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-2 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

arc-agi 77.1% 模型 gemini-3-1-pro · 版本 2 (Verified) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Abstract reasoning · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: ARC-AGI-2 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

aa-intelligence-index 57.3 模型 claude-opus-4-7 · 版本 未说明 · 指标 composite_intelligence_index · 单位 points 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model capabilities section (vega chart) · figure: vega chart; values machine-read from page RSC payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Same chart also lists Gemini 3.1 Pro Preview 57.2, GPT-5.4 xhigh 56.8, Claude Opus 4.6 (max) 53.0, Opus 4.7 non-reasoning 51.8, Opus 4.6 non-reasoning 46.5, GPT-5.4 non-reasoning 35.4.

打开官方来源

mrcr 97.3% 模型 gpt-5-4 · 版本 v2 8-needle 4K-8K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 4K-8K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.

打开官方来源

mrcr 91.4% 模型 gpt-5-4 · 版本 v2 8-needle 8K-16K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 8K-16K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.

打开官方来源

mrcr 97.2% 模型 gpt-5-4 · 版本 v2 8-needle 16K-32K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 16K-32K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.

打开官方来源

mrcr 90.5% 模型 gpt-5-4 · 版本 v2 8-needle 32K-64K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 32K-64K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.

打开官方来源

mrcr 86.0% 模型 gpt-5-4 · 版本 v2 8-needle 64K-128K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 64K-128K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.

打开官方来源

mrcr 79.3% 模型 gpt-5-4 · 版本 v2 8-needle 128K-256K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 128K-256K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.

打开官方来源

mrcr 57.5% 模型 gpt-5-4 · 版本 v2 8-needle 256K-512K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 256K-512K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.

打开官方来源

mrcr 59.2% 模型 claude-opus-4-7 · 版本 v2 8-needle 128K-256K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 128K-256K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.

打开官方来源

mrcr 32.2% 模型 claude-opus-4-7 · 版本 v2 8-needle 512K-1M · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Long context · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: OpenAI MRCR v2 8-needle 512K-1M

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.

打开官方来源

mmmu-pro 81.2% 模型 gpt-5-4 · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - Computer use and vision · table: text table, columns: GPT-5.5 / GPT-5.4 / GPT-5.5 Pro / GPT-5.4 Pro / Claude Opus 4.7 / Gemini 3.1 Pro · row: MMMU Pro (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表(此前仅记录 512K-1M 段)。 Comparison column cited by OpenAI.

打开官方来源