GPT-5.4 / GPT-5.4 Pro / GPT-5.4 mini / GPT-5.4 nano
OpenAI · 2026-03-05 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GPT-5.4
OpenAI 将 GPT-5.4 定位为融合 GPT-5.3-Codex 编码谱系的新旗舰(ChatGPT 内 Thinking 模式,effort none/low/medium/high/xhigh,1M 上下文,原生计算机操作与工具搜索)。已收录 90 余项评测横跨知识工作、编码、计算机操作与长上下文:亮点 GDPval 83.0%、OSWorld-Verified 75.0%。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: GDPval (win or tie)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "xhigh (GPT-5.2 ran at heavy, the lower effort available in ChatGPT)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: "GPT-5.4 sets a new record" at 83.0% vs professionals; prose prints GPT-5.2 as 71.0% while the table prints 70.9% (rounding difference, both recorded). GDPval spans 44 occupations across the 9 largest US-GDP industries; deliverables are real work products.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: SWE-Bench Pro (Public)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "headline; vega latency-cost curve: none 41.1 / low 51.2 / medium 55.3 / high 55.6 / xhigh 57.7",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vega curve also plots GPT-5.3-Codex (none 41.1 / low 51.1 / medium 53.4 / high 55.6 / xhigh 57.2) and GPT-5.2 (none 41.5 / low 47.7 / medium 50.6 / high 53.4 / xhigh 55.6) against latency and tool-call counts.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: OSWorld-Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "headline; vega curve: none 39.8 / low 69.4 / medium 71.8 / high 73.7 / xhigh 75.0",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: record 75.0%, "far above GPT-5.2's 47.3% and above the human performance of 72.4%" (human baseline per footnote 1). First OpenAI general-purpose model with native computer use; operates via Playwright-style code or screenshot-driven mouse/keyboard. Vega curve for GPT-5.2: none 20.8 / low 38.2 / medium 45.4 / high 45.8 / xhigh 47.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "headline; vega curve: none 19.4 / low 41.4 / medium 50.0 / high 50.6 / xhigh 54.6",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}toolathlon id already introduced by a prior batch; reused. Vega curve for GPT-5.2: none 13.0 / low 34.3 / medium 39.5 / high 41.7 / xhigh 45.7, with tool-yield counts; tool yields noted as the better latency metric vs tool calls.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: BrowseComp · figure: vega chart BrowseComp: GPT-5.4 Pro 89.3 / GPT-5.4 82.7 / GPT-5.2 Pro 77.9 / GPT-5.2 65.8; interactive bar/line chart; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: +17 points over GPT-5.2. Protocol footnote: search blocklist excluding answer-bearing sites (broader, updated blacklist for GPT-5.4); tested via the ChatGPT search tool which may differ slightly from API search; internet state evolved between the two models' test dates.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use and vision section · row: WebArena-Verified · quote_snippet: In the WebArena-Verified benchmark [...] GPT-5.4 reached a leading success rate of 67.3%
{
"harness": null,
"tools": "DOM + screenshot-driven interaction combined",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}webarena id exists in data/benchmarks; Verified variant recorded.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use and vision section · row: Online-Mind2Web · quote_snippet: using only screenshot-based observation, GPT-5.4 reached a 92.8% success rate
{
"harness": null,
"tools": "screenshot-only observation",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}online-mind2web id already introduced by a prior batch; reused. Comparison in prose: ChatGPT Atlas agent mode at 70.9%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use and vision section · row: MMMU-Pro · figure: vega chart: GPT-5.4 0.812 vs GPT-5.2 0.795; interactive bar/line chart; values machine-read from page RSC vega-lite payload · quote_snippet: on the MMMU-Pro test, GPT-5.4 reached an 81.2% success rate without tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "xhigh",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmmu-pro not yet migrated into data/benchmarks (introduced in this batch). Vega values cross-check the prose exactly (0.812 / 0.795).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use and vision section · row: OmniDocBench · quote_snippet: on the OmniDocBench test, GPT-5.4 reduced average error (normalized edit distance) to 0.109
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "none (to reflect low-cost, low-latency operation)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}omnidocbench id already introduced by a prior batch; reused. Lower is better. Vega chart cross-checks (0.109 vs 0.140).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FinanceAgent v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OfficeQA
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: officeqa not yet in data/benchmarks. Page label is plain OfficeQA; distinct from the OfficeQA Pro rows printed on the GPT-5.5 page (officeqa-pro).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Terminal-Bench 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: MMMU Pro (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Complements existing no-tools rows (81.2 / 79.5).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: MCP Atlas
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Tau2-bench Telecom
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Tool-use configuration; the same release page also prints a no-tools variant (see rows below).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Frontier Science Research
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 1–3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks BFS 0K–128K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks BFS 256K–1M
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks parents 0–128K (accuracy)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks parents 256K–1M (accuracy)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 4K–8K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 8K–16K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 16K–32K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 32K–64K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 64K–128K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 128K–256K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 256K–512K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 512K–1M
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-1 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-2 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。GPT-5.2 Pro cell prints "54.2% (high)" - the page annotates the effort level in the cell.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation table - no-tools configuration · table: text table, columns: GPT-5.4 (none) / GPT-5.2 (none) / GPT-4.1 · row: Tau2-bench Telecom (none)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
GPT-5.4 Pro
GPT-5.4 Pro 为复杂任务的最大性能档。已收录评测中 Pro 列亮点 GDPval 82.0%、FinanceAgent v1.1 61.5%。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: GPT-5.4 Pro sets the record at 89.3%. Same search-blacklist protocol caveats as the GPT-5.4 row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GDPval
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。GDPval (wins or ties) row; existing rows cover GPT-5.4 / GPT-5.3-Codex / GPT-5.2 from the hero table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FinanceAgent v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Frontier Science Research
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 1–3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-1 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-2 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。GPT-5.2 Pro cell prints "54.2% (high)" - the page annotates the effort level in the cell.
GPT-5.4 mini
GPT-5.4 mini 为更小档位,openai.com 发布 sitemap 中无独立发布页。本次发布未报告其独立评测数值。
- 输入模态
- 官方资料未说明
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。
GPT-5.4 nano
GPT-5.4 nano 为最小档位,openai.com 发布 sitemap 中无独立发布页。本次发布未报告其独立评测数值。
- 输入模态
- 官方资料未说明
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。
未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: GDPval (win or tie)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: GDPval (win or tie)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "heavy",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Table prints 70.9%; prose prints 71.0% for the same comparison (rounding difference).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: SWE-Bench Pro (Public)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: SWE-Bench Pro (Public)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: OSWorld-Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Page footnote: previously reported 64.7%; GPT-5.3-Codex now reaches 74.0% via a new API parameter that preserves original image resolution - an on-page score revision, kept as-is.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: OSWorld-Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: Toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables (knowledge work / coding / tool use sections) · table: text table, columns: GPT-5.4 / GPT-5.3-Codex / GPT-5.2 · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI; vega chart also shows GPT-5.2 Pro at 77.9%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use and vision section · row: WebArena-Verified · quote_snippet: GPT-5.2 was 65.4%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use and vision section · row: Online-Mind2Web · quote_snippet: far above the ChatGPT Atlas agent mode, which achieved 70.9%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}OpenAI's own agent product cited in its release page.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use and vision section · row: MMMU-Pro · quote_snippet: GPT-5.2 was 79.5%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmmu-pro not yet migrated into data/benchmarks (introduced in this batch).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use and vision section · row: OmniDocBench · quote_snippet: a significant improvement over GPT-5.2's 0.140
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Lower is better.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GDPval
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。GDPval (wins or ties) row; existing rows cover GPT-5.4 / GPT-5.3-Codex / GPT-5.2 from the hero table. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FinanceAgent v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FinanceAgent v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OfficeQA
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: officeqa not yet in data/benchmarks. Page label is plain OfficeQA; distinct from the OfficeQA Pro rows printed on the GPT-5.5 page (officeqa-pro). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - knowledge work · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OfficeQA
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。new-benchmark: officeqa not yet in data/benchmarks. Page label is plain OfficeQA; distinct from the OfficeQA Pro rows printed on the GPT-5.5 page (officeqa-pro). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Terminal-Bench 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - coding · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Terminal-Bench 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - computer use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: MMMU Pro (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Complements existing no-tools rows (81.2 / 79.5). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Completes the BrowseComp row (existing rows cover GPT-5.4 / Pro / Codex / 5.2). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: MCP Atlas
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - tool use · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Tau2-bench Telecom
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。Tool-use configuration; the same release page also prints a no-tools variant (see rows below). Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Frontier Science Research
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 1–3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: FrontierMath Tier 4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (no tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - science · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Humanity's Last Exam (with tools)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks BFS 0K–128K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: Graphwalks parents 0–128K (accuracy)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 4K–8K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 8K–16K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 16K–32K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 32K–64K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 64K–128K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - long context · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: OpenAI MRCR v2 8-needle 128K–256K
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-1 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-1 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-2 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。GPT-5.2 Pro cell prints "54.2% (high)" - the page annotates the effort level in the cell. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation tables - abstract reasoning · table: text table, columns: GPT-5.4 / GPT-5.4 Pro / GPT-5.3-Codex / GPT-5.2 / GPT-5.2 Pro · row: ARC-AGI-2 (Verified)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。GPT-5.2 Pro cell prints "54.2% (high)" - the page annotates the effort level in the cell. Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation table - no-tools configuration · table: text table, columns: GPT-5.4 (none) / GPT-5.2 (none) / GPT-4.1 · row: Tau2-bench Telecom (none)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation table - no-tools configuration · table: text table, columns: GPT-5.4 (none) / GPT-5.2 (none) / GPT-4.1 · row: Tau2-bench Telecom (none)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:单元格取自页面评测表。 Comparison column cited by OpenAI.