← 模型目录

OpenAI o3 / OpenAI o4-mini

OpenAI · 2025-04-16 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

OpenAI o3

OpenAI 将 o3 定位为其最强推理模型,在编码、数学、科学与视觉感知等领域推进前沿,并支持以图思考(Thinking with images)。该发布已收录评测覆盖数学、编程竞赛、科学推理、软件工程与智能体工具使用等领域,AIME 2025 配 Python 工具 pass@1 98.4%、100% consensus@8。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

aime24 91.6 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2024 · figure: vega-lite bar chart 'AIME 2024 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2024 Competition Math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Chart values machine-read from the page's embedded vega-lite spec (not OCR).

打开官方来源

aime-25 88.9 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2025 · figure: vega-lite bar chart 'AIME 2025 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2025 Competition Math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark id usage: aime-25 already introduced by the international batch; reused here.

打开官方来源

aime-25 98.4% pass@1, 100% consensus@8 模型 o3 · 版本 with python tool · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models · quote_snippet: o3 with tools allowed performs similarly (98.4% pass@1, 100% consensus@8)

{
  "harness": null,
  "tools": "python interpreter",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1; consensus@8 reported alongside",
  "judge": null
}

Page explicitly warns tool-assisted AIME results must not be compared with no-tools models.

打开官方来源

codeforces 2706 模型 o3 · 版本 未说明 · 指标 codeforces_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: Codeforces · figure: vega-lite bar chart 'Codeforces / Competition Code'; values embedded in page RSC payload · quote_snippet: Codeforces Competition Code

{
  "harness": null,
  "tools": "terminal",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: codeforces already introduced by a prior batch (in the 108-id list); reused here. Prose: o3 'sets new SOTA on Codeforces, SWE-bench (without building a custom model-specific scaffold) and MMMU'.

打开官方来源

gpqa 83.3 模型 o3 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: GPQA Diamond · figure: vega-lite bar chart 'GPQA Diamond / PhD-Level Science Questions'; values embedded in page RSC payload · quote_snippet: GPQA Diamond PhD-Level Science Questions

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Page also contains static image charts o3-o4_mini_GPQA-Pass.png (2352x2300) covering the with-tools GPQA pass configuration; not machine-read.

打开官方来源

hlehle 20.32 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart 'Humanility's Last Exam / Expert-Level Questions Across Subjects'; values embedded in page RSC payload · quote_snippet: Expert-Level Questions Across Subjects

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 24.9 模型 o3 · 版本 with python + browsing tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart; 'o3 (python + browsing** tools)' bar; values embedded in page RSC payload · quote_snippet: o3 (python + browsing tools)

{
  "harness": null,
  "tools": "python + browsing",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

swebench 69.1 模型 o3 · 版本 Verified (fixed subset n=477) · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: SWE-Bench Verified (n=477) · figure: vega-lite bar chart 'SWE-Bench Verified (n=477) / Software Engineering'; values embedded in page RSC payload · quote_snippet: All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks

{
  "harness": "no custom model-specific scaffold (prose claim)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote on page: 'All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks which have been validated on our internal infrastructure.'

打开官方来源

mmmu 82.9 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / multimodal tab (interactive bar chart) · row: MMMU · figure: vega-lite bar chart 'MMMU / College-level visual problem-solving'; values embedded in page RSC payload · quote_snippet: MMMU College-level visual problem-solving

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: o3 'sets new SOTA on Codeforces, SWE-bench (without building a custom model-specific scaffold) and MMMU'.

打开官方来源

mathvista 86.8 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / multimodal tab (interactive bar chart) · row: MathVista · figure: vega-lite bar chart 'MathVista / Visual Math Reasoning'; values embedded in page RSC payload · quote_snippet: MathVista Visual Math Reasoning

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

charxiv-reasoning 78.6 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / multimodal tab (interactive bar chart) · row: CharXiv-Reasoning · figure: vega-lite bar chart 'CharXiv-Reasoning / Scientific Figure Reasoning'; values embedded in page RSC payload · quote_snippet: CharXiv-Reasoning Scientific Figure Reasoning

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

benchmark id charxiv-reasoning already introduced by a prior batch; reused here.

打开官方来源

swe-lancer $86,100 模型 o3 · 版本 IC SWE Diamond · 指标 earned_payout · 单位 usd 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: SWE-Lancer: IC SWE Diamond · figure: vega-lite bar chart 'SWE-Lancer: IC SWE Diamond / Freelance Coding Tasks'; values embedded in page RSC payload · quote_snippet: SWE-Lancer IC SWE Diamond Freelance Coding Tasks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Metric is USD payout earned on freelance tasks, not a percentage. Section footnote: all models evaluated at high reasoning effort (similar to o4-mini-high in ChatGPT).

打开官方来源

multichallenge 56.51 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: Scale MultiChallenge · figure: vega-lite bar chart 'Scale MultiChallenge / Multi-turn instruction following'; values embedded in page RSC payload · quote_snippet: Scale MultiChallenge Multi-turn instruction following

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

benchmark id multichallenge already introduced by a prior batch (Scale MultiChallenge); reused here. Section footnote: all models evaluated at high reasoning effort.

打开官方来源

browsecomp 49.7 模型 o3 · 版本 with python + browsing tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: BrowseComp · figure: vega-lite bar chart 'BrowseComp / Agentic Browsing'; 'o3 with python + browsing*' bar; values embedded in page RSC payload · quote_snippet: BrowseComp Agentic Browsing

{
  "harness": null,
  "tools": "python + browsing",
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

browsecomp id already introduced by a prior batch; reused. Tool-assisted browsing result - not comparable with no-tools models.

打开官方来源

aider 81.3% 模型 o3 · 版本 Polyglot (whole) · 指标 percent_solved · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).

打开官方来源

aider 79.6% 模型 o3 · 版本 Polyglot (diff) · 指标 percent_solved · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).

打开官方来源

tau-bench 52% 模型 o3 · 版本 Airline (high effort) · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.

打开官方来源

tau-bench 70.4% 模型 o3 · 版本 Retail (high effort) · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.

打开官方来源

OpenAI o4-mini

OpenAI 将 o4-mini 定位为面向快速、高性价比推理优化的较小模型,在同等规模与成本下表现突出,同样支持以图思考。该发布已收录评测覆盖数学、编程与智能体工具使用等领域,AIME 2025 配 Python 工具 pass@1 99.5%、100% consensus@8。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

aime24 93.4 模型 o4-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2024 · figure: vega-lite bar chart 'AIME 2024 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2024 Competition Math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: o4-mini 'performs best on AIME 2024 and 2025 benchmarks'.

打开官方来源

aime-25 92.7 模型 o4-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2025 · figure: vega-lite bar chart 'AIME 2025 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2025 Competition Math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

aime-25 99.5% pass@1, 100% consensus@8 模型 o4-mini · 版本 with python tool · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models · quote_snippet: o4-mini achieves a 99.5% pass@1 and 100% consensus@8

{
  "harness": null,
  "tools": "python interpreter",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 8,
  "aggregation": "pass@1; consensus@8 reported alongside",
  "judge": null
}

Page explicitly warns tool-assisted AIME results must not be compared with no-tools models.

打开官方来源

codeforces 2719 模型 o4-mini · 版本 未说明 · 指标 codeforces_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: Codeforces · figure: vega-lite bar chart 'Codeforces / Competition Code'; values embedded in page RSC payload · quote_snippet: Codeforces Competition Code

{
  "harness": null,
  "tools": "terminal",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

o4-mini scores higher than o3 on this chart (2719 vs 2706 Elo).

打开官方来源

gpqa 81.4 模型 o4-mini · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: GPQA Diamond · figure: vega-lite bar chart 'GPQA Diamond / PhD-Level Science Questions'; values embedded in page RSC payload · quote_snippet: GPQA Diamond PhD-Level Science Questions

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 14.28 模型 o4-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: Expert-Level Questions Across Subjects

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 17.7 模型 o4-mini · 版本 with python + browsing tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart; 'o4-mini (with python + browsing** tools)' bar; values embedded in page RSC payload · quote_snippet: o4-mini (with python + browsing tools)

{
  "harness": null,
  "tools": "python + browsing",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

swebench 68.1 模型 o4-mini · 版本 Verified (fixed subset n=477) · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: SWE-Bench Verified (n=477) · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

mmmu 81.6 模型 o4-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / multimodal tab (interactive bar chart) · row: MMMU · figure: vega-lite bar chart 'MMMU / College-level visual problem-solving'; values embedded in page RSC payload · quote_snippet: MMMU College-level visual problem-solving

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

mathvista 84.3 模型 o4-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / multimodal tab (interactive bar chart) · row: MathVista · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: MathVista Visual Math Reasoning

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

charxiv-reasoning 72 模型 o4-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / multimodal tab (interactive bar chart) · row: CharXiv-Reasoning · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: CharXiv-Reasoning Scientific Figure Reasoning

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

swe-lancer $65,792 模型 o4-mini · 版本 IC SWE Diamond · 指标 earned_payout · 单位 usd 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: SWE-Lancer: IC SWE Diamond · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: SWE-Lancer IC SWE Diamond Freelance Coding Tasks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Metric is USD payout earned on freelance tasks, not a percentage.

打开官方来源

multichallenge 42.99 模型 o4-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: Scale MultiChallenge · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: Scale MultiChallenge Multi-turn instruction following

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

browsecomp 28.3 模型 o4-mini · 版本 with python + browsing tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: BrowseComp · figure: vega-lite bar chart; 'o4-mini with python + browsing**' bar; values embedded in page RSC payload · quote_snippet: BrowseComp Agentic Browsing

{
  "harness": null,
  "tools": "python + browsing",
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tool-assisted browsing result - not comparable with no-tools models.

打开官方来源

aider 68.9% 模型 o4-mini · 版本 Polyglot (whole) · 指标 percent_solved · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).

打开官方来源

aider 58.2% 模型 o4-mini · 版本 Polyglot (diff) · 指标 percent_solved · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).

打开官方来源

tau-bench 49.2% 模型 o4-mini · 版本 Airline (high effort) · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.

打开官方来源

tau-bench 65.6% 模型 o4-mini · 版本 Retail (high effort) · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.

打开官方来源

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

aime24 74.3 模型 o1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2024 · figure: vega-lite bar chart 'AIME 2024 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2024 Competition Math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

aime24 87.3 模型 o3-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2024 · figure: vega-lite bar chart 'AIME 2024 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2024 Competition Math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

aime-25 79.2 模型 o1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2025 · figure: vega-lite bar chart 'AIME 2025 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2025 Competition Math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

aime-25 86.5 模型 o3-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: AIME 2025 · figure: vega-lite bar chart 'AIME 2025 / Competition Math'; values embedded in page RSC payload · quote_snippet: AIME 2025 Competition Math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

codeforces 1891 模型 o1 · 版本 未说明 · 指标 codeforces_elo · 单位 elo 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: Codeforces · figure: vega-lite bar chart 'Codeforces / Competition Code'; values embedded in page RSC payload · quote_snippet: Codeforces Competition Code

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

codeforces 2073 模型 o3-mini · 版本 未说明 · 指标 codeforces_elo · 单位 elo 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: Codeforces · figure: vega-lite bar chart 'Codeforces / Competition Code'; values embedded in page RSC payload · quote_snippet: Codeforces Competition Code

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

gpqa 78 模型 o1 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: GPQA Diamond · figure: vega-lite bar chart 'GPQA Diamond / PhD-Level Science Questions'; values embedded in page RSC payload · quote_snippet: GPQA Diamond PhD-Level Science Questions

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

gpqa 77 模型 o3-mini · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: GPQA Diamond · figure: vega-lite bar chart 'GPQA Diamond / PhD-Level Science Questions'; values embedded in page RSC payload · quote_snippet: GPQA Diamond PhD-Level Science Questions

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

hlehle 26.6 模型 openai-deep-research · 版本 python + browser tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart; 'Deep research' bar; values embedded in page RSC payload · quote_snippet: Deep research

{
  "harness": null,
  "tools": "python + browser",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

OpenAI's own agent product (not an o-series model) included in the comparison chart.

打开官方来源

hlehle 13.4 模型 o3-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: Humanity's Last Exam · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: Expert-Level Questions Across Subjects

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

swebench 48.9 模型 o1 · 版本 Verified (fixed subset n=477) · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: SWE-Bench Verified (n=477) · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

swebench 49.3 模型 o3-mini · 版本 Verified (fixed subset n=477) · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / coding tab (interactive bar chart) · row: SWE-Bench Verified (n=477) · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

mmmu 77.6 模型 o1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / multimodal tab (interactive bar chart) · row: MMMU · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: MMMU College-level visual problem-solving

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

mathvista 71.8 模型 o1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / multimodal tab (interactive bar chart) · row: MathVista · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: MathVista Visual Math Reasoning

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

charxiv-reasoning 55.1 模型 o1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: What's new in our models / multimodal tab (interactive bar chart) · row: CharXiv-Reasoning · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: CharXiv-Reasoning Scientific Figure Reasoning

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

swe-lancer $43,958 模型 o1 · 版本 IC SWE Diamond · 指标 earned_payout · 单位 usd 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: SWE-Lancer: IC SWE Diamond · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: SWE-Lancer IC SWE Diamond Freelance Coding Tasks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

swe-lancer $33,833 模型 o3-mini · 版本 IC SWE Diamond · 指标 earned_payout · 单位 usd 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: SWE-Lancer: IC SWE Diamond · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: SWE-Lancer IC SWE Diamond Freelance Coding Tasks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

multichallenge 44.93 模型 o1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: Scale MultiChallenge · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: Scale MultiChallenge Multi-turn instruction following

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

multichallenge 39.89 模型 o3-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: Scale MultiChallenge · figure: vega-lite bar chart; values embedded in page RSC payload · quote_snippet: Scale MultiChallenge Multi-turn instruction following

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

browsecomp 1.9 模型 gpt-4o · 版本 with browsing tool · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: BrowseComp · figure: vega-lite bar chart; '4o + browsing' bar; values embedded in page RSC payload · quote_snippet: 4o + browsing

{
  "harness": null,
  "tools": "browsing",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation model score cited by OpenAI in its own release chart.

打开官方来源

browsecomp 51.5 模型 openai-deep-research · 版本 deep research agent · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Instruction following and agentic tool use (interactive bar chart) · row: BrowseComp · figure: vega-lite bar chart; 'Deep research' bar; values embedded in page RSC payload · quote_snippet: Deep research

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

OpenAI's own agent product (not an o-series model) included in the comparison chart.

打开官方来源

aider 64.4% 模型 o1 · 版本 Polyglot (whole) · 指标 percent_solved · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).

打开官方来源

aider 61.7% 模型 o1 · 版本 Polyglot (diff) · 指标 percent_solved · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).

打开官方来源

aider 66.7% 模型 o3-mini · 版本 Polyglot (whole) · 指标 percent_solved · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).

打开官方来源

aider 60.4% 模型 o3-mini · 版本 Polyglot (diff) · 指标 percent_solved · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - code editing · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.644 for whole-edit format); bars compared at high reasoning effort (chart model labels o1-high / o3-mini-high / o3-high / o4-mini-high).

打开官方来源

tau-bench 50% 模型 o1 · 版本 Airline (high effort) · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.

打开官方来源

tau-bench 70.8% 模型 o1 · 版本 Retail (high effort) · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.

打开官方来源

tau-bench 32.4% 模型 o3-mini · 版本 Airline (high effort) · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.

打开官方来源

tau-bench 57.6% 模型 o3-mini · 版本 Retail (high effort) · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart data prints the rate as a fraction (e.g. 0.5); chart model labels carry -high effort suffix. The chart dataset also holds a third color-series pair of bars for o3-high / o4-mini-high Retail (0.035 / 0.062) whose series name is not printed in the payload; not recorded as rows.

打开官方来源

hlehle 8.12% 模型 o1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - expert-level questions · figure: vega chart "Humanity's Last Exam" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Completes the HLE chart row (o1-pro bar was previously unrecorded).

打开官方来源