OpenAI o3 / OpenAI o4-mini
OpenAI · 2025-04-16 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
OpenAI o3
OpenAI 将 o3 定位为其最强推理模型,在编码、数学、科学与视觉感知等领域推进前沿,并支持以图思考(Thinking with images)。该发布已收录评测覆盖数学、编程竞赛、科学推理、软件工程与智能体工具使用等领域,AIME 2025 配 Python 工具 pass@1 98.4%、100% consensus@8。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2024 · figure: vega-lite bar chart 'AIME 2024 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2024 Competition Math
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Chart values machine-read from the page's embedded vega-lite spec (not OCR).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2025 · figure: vega-lite bar chart 'AIME 2025 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2025 Competition Math
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark id usage: aime-25 already introduced by the international batch; reused here.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models · quote_snippet: o3 with tools allowed performs similarly (98.4% pass@1, 100% consensus@8)
{
"harness": null,
"tools": "python interpreter",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1; consensus@8 reported alongside",
"judge": null
}Page explicitly warns tool-assisted AIME results must not be compared with no-tools models.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: Codeforces · figure: vega-lite bar chart 'Codeforces / Competition Code'; values embedded in page RSC payload · quote_snippet: Codeforces Competition Code
{
"harness": null,
"tools": "terminal",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: codeforces already introduced by a prior batch (in the 108-id list); reused here. Prose: o3 'sets new SOTA on Codeforces, SWE-bench (without building a custom model-specific scaffold) and MMMU'.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: GPQA Diamond · figure: vega-lite bar chart 'GPQA Diamond / PhD-Level Science Questions'; values embedded in page RSC payload · quote_snippet: GPQA Diamond PhD-Level Science Questions
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Page also contains static image charts o3-o4_mini_GPQA-Pass.png (2352x2300) covering the with-tools GPQA pass configuration; not machine-read.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart 'Humanility's Last Exam / Expert-Level Questions Across Subjects'; values embedded in page RSC payload · quote_snippet: Expert-Level Questions Across Subjects
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart; 'o3 (python + browsing** tools)' bar; values embedded in page RSC payload · quote_snippet: o3 (python + browsing tools)
{
"harness": null,
"tools": "python + browsing",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: SWE-Bench Verified (n=477) · figure: vega-lite bar chart 'SWE-Bench Verified (n=477) / Software Engineering'; values embedded in page RSC payload · quote_snippet: All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks
{
"harness": "no custom model-specific scaffold (prose claim)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote on page: 'All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks which have been validated on our internal infrastructure.'
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / multimodal tab (interactive bar chart) · row: MMMU · figure: vega-lite bar chart 'MMMU / College-level visual problem-solving'; values embedded in page RSC payload · quote_snippet: MMMU College-level visual problem-solving
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: o3 'sets new SOTA on Codeforces, SWE-bench (without building a custom model-specific scaffold) and MMMU'.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / multimodal tab (interactive bar chart) · row: MathVista · figure: vega-lite bar chart 'MathVista / Visual Math Reasoning'; values embedded in page RSC payload · quote_snippet: MathVista Visual Math Reasoning
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / multimodal tab (interactive bar chart) · row: CharXiv-Reasoning · figure: vega-lite bar chart 'CharXiv-Reasoning / Scientific Figure Reasoning'; values embedded in page RSC payload · quote_snippet: CharXiv-Reasoning Scientific Figure Reasoning
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}benchmark id charxiv-reasoning already introduced by a prior batch; reused here.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: SWE-Lancer: IC SWE Diamond · figure: vega-lite bar chart 'SWE-Lancer: IC SWE Diamond / Freelance Coding Tasks'; values embedded in page RSC payload · quote_snippet: SWE-Lancer IC SWE Diamond Freelance Coding Tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Metric is USD payout earned on freelance tasks, not a percentage. Section footnote: all models evaluated at high reasoning effort (similar to o4-mini-high in ChatGPT).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: Scale MultiChallenge · figure: vega-lite bar chart 'Scale MultiChallenge / Multi-turn instruction following'; values embedded in page RSC payload · quote_snippet: Scale MultiChallenge Multi-turn instruction following
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}benchmark id multichallenge already introduced by a prior batch (Scale MultiChallenge); reused here. Section footnote: all models evaluated at high reasoning effort.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: BrowseComp · figure: vega-lite bar chart 'BrowseComp / Agentic Browsing'; 'o3 with python + browsing*' bar; values embedded in page RSC payload · quote_snippet: BrowseComp Agentic Browsing
{
"harness": null,
"tools": "python + browsing",
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}browsecomp id already introduced by a prior batch; reused. Tool-assisted browsing result - not comparable with no-tools models.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.
OpenAI o4-mini
OpenAI 将 o4-mini 定位为面向快速、高性价比推理优化的较小模型,在同等规模与成本下表现突出,同样支持以图思考。该发布已收录评测覆盖数学、编程与智能体工具使用等领域,AIME 2025 配 Python 工具 pass@1 99.5%、100% consensus@8。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2024 · figure: vega-lite bar chart 'AIME 2024 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2024 Competition Math
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: o4-mini 'performs best on AIME 2024 and 2025 benchmarks'.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2025 · figure: vega-lite bar chart 'AIME 2025 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2025 Competition Math
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models · quote_snippet: o4-mini achieves a 99.5% pass@1 and 100% consensus@8
{
"harness": null,
"tools": "python interpreter",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 8,
"aggregation": "pass@1; consensus@8 reported alongside",
"judge": null
}Page explicitly warns tool-assisted AIME results must not be compared with no-tools models.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: Codeforces · figure: vega-lite bar chart 'Codeforces / Competition Code'; values embedded in page RSC payload · quote_snippet: Codeforces Competition Code
{
"harness": null,
"tools": "terminal",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}o4-mini scores higher than o3 on this chart (2719 vs 2706 Elo).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: GPQA Diamond · figure: vega-lite bar chart 'GPQA Diamond / PhD-Level Science Questions'; values embedded in page RSC payload · quote_snippet: GPQA Diamond PhD-Level Science Questions
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: Expert-Level Questions Across Subjects
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart; 'o4-mini (with python + browsing** tools)' bar; values embedded in page RSC payload · quote_snippet: o4-mini (with python + browsing tools)
{
"harness": null,
"tools": "python + browsing",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: SWE-Bench Verified (n=477) · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / multimodal tab (interactive bar chart) · row: MMMU · figure: vega-lite bar chart 'MMMU / College-level visual problem-solving'; values embedded in page RSC payload · quote_snippet: MMMU College-level visual problem-solving
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / multimodal tab (interactive bar chart) · row: MathVista · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: MathVista Visual Math Reasoning
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / multimodal tab (interactive bar chart) · row: CharXiv-Reasoning · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: CharXiv-Reasoning Scientific Figure Reasoning
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: SWE-Lancer: IC SWE Diamond · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: SWE-Lancer IC SWE Diamond Freelance Coding Tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Metric is USD payout earned on freelance tasks, not a percentage.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: Scale MultiChallenge · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: Scale MultiChallenge Multi-turn instruction following
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: BrowseComp · figure: vega-lite bar chart; 'o4-mini with python + browsing**' bar; values embedded in page RSC payload · quote_snippet: BrowseComp Agentic Browsing
{
"harness": null,
"tools": "python + browsing",
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Tool-assisted browsing result - not comparable with no-tools models.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.
未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2024 · figure: vega-lite bar chart 'AIME 2024 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2024 Competition Math
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2024 · figure: vega-lite bar chart 'AIME 2024 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2024 Competition Math
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2025 · figure: vega-lite bar chart 'AIME 2025 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2025 Competition Math
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2025 · figure: vega-lite bar chart 'AIME 2025 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2025 Competition Math
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: Codeforces · figure: vega-lite bar chart 'Codeforces / Competition Code'; values embedded in page RSC payload · quote_snippet: Codeforces Competition Code
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: Codeforces · figure: vega-lite bar chart 'Codeforces / Competition Code'; values embedded in page RSC payload · quote_snippet: Codeforces Competition Code
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: GPQA Diamond · figure: vega-lite bar chart 'GPQA Diamond / PhD-Level Science Questions'; values embedded in page RSC payload · quote_snippet: GPQA Diamond PhD-Level Science Questions
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: GPQA Diamond · figure: vega-lite bar chart 'GPQA Diamond / PhD-Level Science Questions'; values embedded in page RSC payload · quote_snippet: GPQA Diamond PhD-Level Science Questions
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart; 'Deep research' bar; values embedded in page RSC payload · quote_snippet: Deep research
{
"harness": null,
"tools": "python + browser",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}OpenAI's own agent product (not an o-series model) included in the comparison chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: Expert-Level Questions Across Subjects
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: SWE-Bench Verified (n=477) · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / coding tab (interactive bar chart) · row: SWE-Bench Verified (n=477) · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / multimodal tab (interactive bar chart) · row: MMMU · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: MMMU College-level visual problem-solving
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / multimodal tab (interactive bar chart) · row: MathVista · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: MathVista Visual Math Reasoning
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: What's new in our models / multimodal tab (interactive bar chart) · row: CharXiv-Reasoning · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: CharXiv-Reasoning Scientific Figure Reasoning
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: SWE-Lancer: IC SWE Diamond · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: SWE-Lancer IC SWE Diamond Freelance Coding Tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: SWE-Lancer: IC SWE Diamond · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: SWE-Lancer IC SWE Diamond Freelance Coding Tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: Scale MultiChallenge · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: Scale MultiChallenge Multi-turn instruction following
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: Scale MultiChallenge · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: Scale MultiChallenge Multi-turn instruction following
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: BrowseComp · figure: vega-lite bar chart; '4o + browsing' bar; values embedded in page RSC payload · quote_snippet: 4o + browsing
{
"harness": null,
"tools": "browsing",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prior-generation model score cited by OpenAI in its own release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Instruction following and agentic tool use (interactive bar chart) · row: BrowseComp · figure: vega-lite bar chart; 'Deep research' bar; values embedded in page RSC payload · quote_snippet: Deep research
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}OpenAI's own agent product (not an o-series model) included in the comparison chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - expert-level questions · figure: vega chart "Humanity's Last Exam" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Completes the HLE chart row (o1-pro bar was previously unrecorded).