← 模型目录

GPT-5.4 / GPT-5.4 Pro / GPT-5.4 mini / GPT-5.4 nano

OpenAI · 2026-03-05 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GPT-5.4

OpenAI 将 GPT-5.4 定位为融合 GPT-5.3-Codex 编码谱系的新旗舰(ChatGPT 内 Thinking 模式,effort none/low/medium/high/xhigh,1M 上下文,原生计算机操作与工具搜索)。已收录 90 余项评测横跨知识工作、编码、计算机操作与长上下文:亮点 GDPval 83.0%、OSWorld-Verified 75.0%。

输入模态
文本 / 图像
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

gdpval 83.0% 模型 gpt-5-4 · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: GDPval (win or tie)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "xhigh (GPT-5.2 ran at heavy, the lower effort available in ChatGPT)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: "GPT-5.4 sets a new record" at 83.0% vs professionals; prose prints GPT-5.2 as 71.0% while the table prints 70.9% (rounding difference, both recorded). GDPval spans 44 occupations across the 9 largest US-GDP industries; deliverables are real work products.

打开官方来源

swebench-pro 57.7% 模型 gpt-5-4 · 版本 Public · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: SWE-Bench Pro (Public)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "headline; vega latency-cost curve: none 41.1 / low 51.2 / medium 55.3 / high 55.6 / xhigh 57.7",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vega curve also plots GPT-5.3-Codex (none 41.1 / low 51.1 / medium 53.4 / high 55.6 / xhigh 57.2) and GPT-5.2 (none 41.5 / low 47.7 / medium 50.6 / high 53.4 / xhigh 55.6) against latency and tool-call counts.

打开官方来源

osworld 75.0% 模型 gpt-5-4 · 版本 Verified · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: OSWorld-Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "headline; vega curve: none 39.8 / low 69.4 / medium 71.8 / high 73.7 / xhigh 75.0",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: record 75.0%, "far above GPT-5.2's 47.3% and above the human performance of 72.4%" (human baseline per footnote 1). First OpenAI general-purpose model with native computer use; operates via Playwright-style code or screenshot-driven mouse/keyboard. Vega curve for GPT-5.2: none 20.8 / low 38.2 / medium 45.4 / high 45.8 / xhigh 47.3.

打开官方来源

toolathlon 54.6% 模型 gpt-5-4 · 版本 未说明 · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "headline; vega curve: none 19.4 / low 41.4 / medium 50.0 / high 50.6 / xhigh 54.6",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

toolathlon id already introduced by a prior batch; reused. Vega curve for GPT-5.2: none 13.0 / low 34.3 / medium 39.5 / high 41.7 / xhigh 45.7, with tool-yield counts; tool yields noted as the better latency metric vs tool calls.

打开官方来源

browsecomp 82.7% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: BrowseComp · figure: vega chart BrowseComp: GPT-5.4 Pro 89.3 / GPT-5.4 82.7 / GPT-5.2 Pro 77.9 / GPT-5.2 65.8; interactive bar/line chart; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: +17 points over GPT-5.2. Protocol footnote: search blocklist excluding answer-bearing sites (broader, updated blacklist for GPT-5.4); tested via the ChatGPT search tool which may differ slightly from API search; internet state evolved between the two models' test dates.

打开官方来源

webarena 67.3% 模型 gpt-5-4 · 版本 Verified · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Computer use and vision section · row: WebArena-Verified · quote_snippet: In the WebArena-Verified benchmark [...] GPT-5.4 reached a leading success rate of 67.3%

{
  "harness": null,
  "tools": "DOM + screenshot-driven interaction combined",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

webarena id exists in data/benchmarks; Verified variant recorded.

打开官方来源

online-mind2web 92.8% 模型 gpt-5-4 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Computer use and vision section · row: Online-Mind2Web · quote_snippet: using only screenshot-based observation, GPT-5.4 reached a 92.8% success rate

{
  "harness": null,
  "tools": "screenshot-only observation",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

online-mind2web id already introduced by a prior batch; reused. Comparison in prose: ChatGPT Atlas agent mode at 70.9%.

打开官方来源

mmmu-pro 81.2% 模型 gpt-5-4 · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Computer use and vision section · row: MMMU-Pro · figure: vega chart: GPT-5.4 0.812 vs GPT-5.2 0.795; interactive bar/line chart; values machine-read from page RSC vega-lite payload · quote_snippet: on the MMMU-Pro test, GPT-5.4 reached an 81.2% success rate without tools

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "xhigh",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmmu-pro not yet migrated into data/benchmarks (introduced in this batch). Vega values cross-check the prose exactly (0.812 / 0.795).

打开官方来源

omnidocbench 0.109 模型 gpt-5-4 · 版本 未说明 · 指标 avg_normalized_edit_distance · 单位 normalized_edit_distance 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Computer use and vision section · row: OmniDocBench · quote_snippet: on the OmniDocBench test, GPT-5.4 reduced average error (normalized edit distance) to 0.109

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "none (to reflect low-cost, low-latency operation)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

omnidocbench id already introduced by a prior batch; reused. Lower is better. Vega chart cross-checks (0.109 vs 0.140).

打开官方来源

finance-agent 56.0% 模型 gpt-5-4 · 版本 v1.1 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FinanceAgent v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

officeqa 68.1% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OfficeQA

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: officeqa not yet in data/benchmarks. Page label is plain OfficeQA; distinct from the OfficeQA Pro rows printed on the GPT-5.5 page (officeqa-pro).

打开官方来源

terminalbench 75.1% 模型 gpt-5-4 · 版本 2.0 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Terminal-Bench 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mmmu-pro 82.1% 模型 gpt-5-4 · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: MMMU Pro (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Complements existing no-tools rows (81.2 / 79.5).

打开官方来源

mcp-atlas 67.2% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: MCP Atlas

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

tau2-bench 98.9% 模型 gpt-5-4 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Tau2-bench Telecom

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Tool-use configuration; the same release page also prints a no-tools variant (see rows below).

打开官方来源

frontier-science-research 33.0% 模型 gpt-5-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Frontier Science Research

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

frontiermath 47.6% 模型 gpt-5-4 · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 1–3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

frontiermath 27.1% 模型 gpt-5-4 · 版本 Tier 4 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

gpqa 92.8% 模型 gpt-5-4 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

hlehle 39.8% 模型 gpt-5-4 · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

hlehle 52.1% 模型 gpt-5-4 · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

graphwalks 93.0% 模型 gpt-5-4 · 版本 BFS 0K–128K · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks BFS 0K–128K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

graphwalks 21.4% 模型 gpt-5-4 · 版本 BFS 256K–1M · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks BFS 256K–1M

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

graphwalks 89.8% 模型 gpt-5-4 · 版本 parents 0–128K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks parents 0–128K (accuracy)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

graphwalks 32.4% 模型 gpt-5-4 · 版本 parents 256K–1M · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks parents 256K–1M (accuracy)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 97.3% 模型 gpt-5-4 · 版本 v2 8-needle 4K–8K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 4K–8K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 91.4% 模型 gpt-5-4 · 版本 v2 8-needle 8K–16K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 8K–16K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 97.2% 模型 gpt-5-4 · 版本 v2 8-needle 16K–32K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 16K–32K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 90.5% 模型 gpt-5-4 · 版本 v2 8-needle 32K–64K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 32K–64K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 86.0% 模型 gpt-5-4 · 版本 v2 8-needle 64K–128K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 64K–128K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 79.3% 模型 gpt-5-4 · 版本 v2 8-needle 128K–256K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 128K–256K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 57.5% 模型 gpt-5-4 · 版本 v2 8-needle 256K–512K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 256K–512K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

mrcr 36.6% 模型 gpt-5-4 · 版本 v2 8-needle 512K–1M · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 512K–1M

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

arc-agi 93.7% 模型 gpt-5-4 · 版本 1 (Verified) · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-1 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

arc-agi 73.3% 模型 gpt-5-4 · 版本 2 (Verified) · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-2 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。GPT-5.2 Pro cell prints "54.2% (high)" - the page annotates the effort level in the cell.

打开官方来源

tau2-bench 64.3% 模型 gpt-5-4 · 版本 no tools · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation table - no-tools configuration · table: text table, columns: GPT-5.4 (none) / GPT-5.2 (none) / GPT-4.1 · row: Tau2-bench Telecom (none)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

GPT-5.4 Pro

GPT-5.4 Pro 为复杂任务的最大性能档。已收录评测中 Pro 列亮点 GDPval 82.0%、FinanceAgent v1.1 61.5%。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

browsecomp 89.3% 模型 gpt-5-4-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: GPT-5.4 Pro sets the record at 89.3%. Same search-blacklist protocol caveats as the GPT-5.4 row.

打开官方来源

gdpval 82.0% 模型 gpt-5-4-pro · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GDPval

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。GDPval (wins or ties) row; existing rows cover GPT-5.4 / GPT-5.3-Codex / GPT-5.2 from the hero table.

打开官方来源

finance-agent 61.5% 模型 gpt-5-4-pro · 版本 v1.1 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FinanceAgent v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

frontier-science-research 36.7% 模型 gpt-5-4-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Frontier Science Research

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

frontiermath 50.0% 模型 gpt-5-4-pro · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 1–3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

frontiermath 38.0% 模型 gpt-5-4-pro · 版本 Tier 4 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

gpqa 94.4% 模型 gpt-5-4-pro · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

hlehle 42.7% 模型 gpt-5-4-pro · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

hlehle 58.7% 模型 gpt-5-4-pro · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

arc-agi 94.5% 模型 gpt-5-4-pro · 版本 1 (Verified) · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-1 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。

打开官方来源

arc-agi 83.3% 模型 gpt-5-4-pro · 版本 2 (Verified) · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-2 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。GPT-5.2 Pro cell prints "54.2% (high)" - the page annotates the effort level in the cell.

打开官方来源

GPT-5.4 mini

GPT-5.4 mini 为更小档位,openai.com 发布 sitemap 中无独立发布页。本次发布未报告其独立评测数值。

输入模态
官方资料未说明
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。

GPT-5.4 nano

GPT-5.4 nano 为最小档位,openai.com 发布 sitemap 中无独立发布页。本次发布未报告其独立评测数值。

输入模态
官方资料未说明
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

gdpval 70.9% 模型 gpt-5-3-codex · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: GDPval (win or tie)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI.

打开官方来源

gdpval 70.9% 模型 gpt-5-2 · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: GDPval (win or tie)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "heavy",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Table prints 70.9%; prose prints 71.0% for the same comparison (rounding difference).

打开官方来源

swebench-pro 56.8% 模型 gpt-5-3-codex · 版本 Public · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: SWE-Bench Pro (Public)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI.

打开官方来源

swebench-pro 55.6% 模型 gpt-5-2 · 版本 Public · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: SWE-Bench Pro (Public)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI.

打开官方来源

osworld 74.0%* 模型 gpt-5-3-codex · 版本 Verified · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: OSWorld-Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Page footnote: previously reported 64.7%; GPT-5.3-Codex now reaches 74.0% via a new API parameter that preserves original image resolution - an on-page score revision, kept as-is.

打开官方来源

osworld 47.3% 模型 gpt-5-2 · 版本 Verified · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: OSWorld-Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI.

打开官方来源

toolathlon 51.9% 模型 gpt-5-3-codex · 版本 未说明 · 指标 pass_at_1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI.

打开官方来源

toolathlon 46.3% 模型 gpt-5-2 · 版本 未说明 · 指标 pass_at_1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: Toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI.

打开官方来源

browsecomp 77.3% 模型 gpt-5-3-codex · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI.

打开官方来源

browsecomp 65.8% 模型 gpt-5-2 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI; vega chart also shows GPT-5.2 Pro at 77.9%.

打开官方来源

webarena 65.4% 模型 gpt-5-2 · 版本 Verified · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Computer use and vision section · row: WebArena-Verified · quote_snippet: GPT-5.2 was 65.4%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

online-mind2web 70.9% 模型 chatgpt-atlas-agent · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Computer use and vision section · row: Online-Mind2Web · quote_snippet: far above the ChatGPT Atlas agent mode, which achieved 70.9%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

OpenAI's own agent product cited in its release page.

打开官方来源

mmmu-pro 79.5% 模型 gpt-5-2 · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Computer use and vision section · row: MMMU-Pro · quote_snippet: GPT-5.2 was 79.5%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmmu-pro not yet migrated into data/benchmarks (introduced in this batch).

打开官方来源

omnidocbench 0.140 模型 gpt-5-2 · 版本 未说明 · 指标 avg_normalized_edit_distance · 单位 normalized_edit_distance 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Computer use and vision section · row: OmniDocBench · quote_snippet: a significant improvement over GPT-5.2's 0.140

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Lower is better.

打开官方来源

gdpval 74.1% 模型 gpt-5-2-pro · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GDPval

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。GDPval (wins or ties) row; existing rows cover GPT-5.4 / GPT-5.3-Codex / GPT-5.2 from the hero table. Comparison column cited by OpenAI.

打开官方来源

finance-agent 54.0% 模型 gpt-5-3-codex · 版本 v1.1 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FinanceAgent v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

finance-agent 59.5% 模型 gpt-5-2 · 版本 v1.1 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FinanceAgent v1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

officeqa 65.1% 模型 gpt-5-3-codex · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OfficeQA

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: officeqa not yet in data/benchmarks. Page label is plain OfficeQA; distinct from the OfficeQA Pro rows printed on the GPT-5.5 page (officeqa-pro). Comparison column cited by OpenAI.

打开官方来源

officeqa 63.1% 模型 gpt-5-2 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OfficeQA

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: officeqa not yet in data/benchmarks. Page label is plain OfficeQA; distinct from the OfficeQA Pro rows printed on the GPT-5.5 page (officeqa-pro). Comparison column cited by OpenAI.

打开官方来源

terminalbench 77.3% 模型 gpt-5-3-codex · 版本 2.0 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Terminal-Bench 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

terminalbench 62.2% 模型 gpt-5-2 · 版本 2.0 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - coding · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Terminal-Bench 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mmmu-pro 80.4% 模型 gpt-5-2 · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - computer use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: MMMU Pro (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Complements existing no-tools rows (81.2 / 79.5). Comparison column cited by OpenAI.

打开官方来源

browsecomp 77.9% 模型 gpt-5-2-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Completes the BrowseComp row (existing rows cover GPT-5.4 / Pro / Codex / 5.2). Comparison column cited by OpenAI.

打开官方来源

mcp-atlas 60.6% 模型 gpt-5-2 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: MCP Atlas

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

tau2-bench 98.7% 模型 gpt-5-2 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - tool use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Tau2-bench Telecom

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。Tool-use configuration; the same release page also prints a no-tools variant (see rows below). Comparison column cited by OpenAI.

打开官方来源

frontier-science-research 25.2% 模型 gpt-5-2 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Frontier Science Research

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

frontiermath 40.7% 模型 gpt-5-2 · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 1–3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

frontiermath 18.8% 模型 gpt-5-2 · 版本 Tier 4 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

frontiermath 31.3% 模型 gpt-5-2-pro · 版本 Tier 4 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gpqa 92.6% 模型 gpt-5-3-codex · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gpqa 92.4% 模型 gpt-5-2 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

gpqa 93.2% 模型 gpt-5-2-pro · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

hlehle 34.5% 模型 gpt-5-2 · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

hlehle 36.6% 模型 gpt-5-2-pro · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (no tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

hlehle 45.5% 模型 gpt-5-2 · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

hlehle 50.0% 模型 gpt-5-2-pro · 版本 with tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (with tools)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

graphwalks 94.0% 模型 gpt-5-2 · 版本 BFS 0K–128K · 指标 f1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks BFS 0K–128K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

graphwalks 89.0% 模型 gpt-5-2 · 版本 parents 0–128K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks parents 0–128K (accuracy)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mrcr 98.2% 模型 gpt-5-2 · 版本 v2 8-needle 4K–8K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 4K–8K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mrcr 89.3% 模型 gpt-5-2 · 版本 v2 8-needle 8K–16K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 8K–16K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mrcr 95.3% 模型 gpt-5-2 · 版本 v2 8-needle 16K–32K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 16K–32K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mrcr 92.0% 模型 gpt-5-2 · 版本 v2 8-needle 32K–64K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 32K–64K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mrcr 85.6% 模型 gpt-5-2 · 版本 v2 8-needle 64K–128K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 64K–128K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

mrcr 77.0% 模型 gpt-5-2 · 版本 v2 8-needle 128K–256K · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 128K–256K

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

arc-agi 86.2% 模型 gpt-5-2 · 版本 1 (Verified) · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-1 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

arc-agi 90.5% 模型 gpt-5-2-pro · 版本 1 (Verified) · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-1 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

arc-agi 52.9% 模型 gpt-5-2 · 版本 2 (Verified) · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-2 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。GPT-5.2 Pro cell prints "54.2% (high)" - the page annotates the effort level in the cell. Comparison column cited by OpenAI.

打开官方来源

arc-agi 54.2% (high) 模型 gpt-5-2-pro · 版本 2 (Verified) · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-2 (Verified)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。GPT-5.2 Pro cell prints "54.2% (high)" - the page annotates the effort level in the cell. Comparison column cited by OpenAI.

打开官方来源

tau2-bench 43.6% 模型 gpt-4-1 · 版本 no tools · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation table - no-tools configuration · table: text table, columns: GPT-5.4 (none) / GPT-5.2 (none) / GPT-4.1 · row: Tau2-bench Telecom (none)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源

tau2-bench 57.2% 模型 gpt-5-2 · 版本 no tools · 指标 pass_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation table - no-tools configuration · table: text table, columns: GPT-5.4 (none) / GPT-5.2 (none) / GPT-4.1 · row: Tau2-bench Telecom (none)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.

打开官方来源