← 模型目录

GPT-5 / GPT-5 Pro

OpenAI · 2025-08-07 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GPT-5

GPT-5 是 OpenAI 的统一思考模型,发布文称其在 AIME 2025(无工具 94.6%)、SWE-bench Verified 74.9%、HealthBench Hard 46.2% 等刷新 SOTA,并大幅降低事实错误与谄媚率。评测覆盖竞赛数学、编码、指令跟随与代理、多模态与健康:GPQA Diamond 88.4%、MMMU 84.2%、BrowseComp 54.9%。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

aime-25 94.6 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose anchor: "94.6% on AIME 2025 without tools". Stacked bar (thinking 32.7 + non-thinking 61.9 = 94.6); chart labelValue 94.6.

打开官方来源

aime-25 99.6 模型 gpt-5 · 版本 with python tool · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote on page: tool-assisted AIME results must not be compared with no-tools models.

打开官方来源

frontiermath 26.3 模型 gpt-5 · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 13.5 模型 gpt-5 · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hmmt25 96.7 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: HMMT · figure: interactive bar chart 'HMMT'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hmmt25 93.3 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: HMMT · figure: interactive bar chart 'HMMT'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

gpqa 87.3 模型 gpt-5 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

gpqa 85.7 模型 gpt-5 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Stacked bar (thinking 7.9 + non-thinking 77.8 = 85.7); chart labelValue 85.7.

打开官方来源

hlehle 24.8 模型 gpt-5 · 版本 Full Set · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Stacked bar (thinking 18.5 + non-thinking 6.3 = 24.8); chart labelValue 24.8.

打开官方来源

swebench 74.9 模型 gpt-5 · 版本 Verified (fixed subset n=477) · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: SWE-bench Verified (n=477) · figure: interactive bar chart 'SWE-bench Verified (n=477)'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high (default); effort line chart: low 69.1 / medium 72.4 / high 74.9",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: "74.9% on SWE-bench Verified". Companion line chart (tokens vs accuracy) gives GPT-5 low/medium/high 69.1/72.4/74.9 and o3 low/medium/high 63.8/67.1/69.1 at 4202-13741 output tokens. Footnote: fixed subset of n=477 verified tasks validated on OpenAI internal infrastructure.

打开官方来源

swebench 52.8 模型 gpt-5 · 版本 Verified (fixed subset n=477) · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: SWE-bench Verified (n=477) · figure: interactive bar chart 'SWE-bench Verified (n=477)'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

aider 88 模型 gpt-5 · 版本 Polyglot · 指标 percent_solved · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Aider Polyglot · figure: interactive bar chart 'Aider Polyglot'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: "88% on Aider-Polyglot". Stacked bar (thinking 61.3 + non-thinking 26.7 = 88); chart labelValue 88.

打开官方来源

multichallenge 69.6 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Scale MultiChallenge · figure: interactive bar chart 'Scale MultiChallenge'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Stacked bar (thinking 14.7 + non-thinking 54.9 = 69.6); chart labelValue 69.6.

打开官方来源

browsecomp 54.9 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: BrowseComp · figure: interactive bar chart 'BrowseComp'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

mmmu 84.2 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: MMMU · figure: interactive bar chart 'MMMU'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: "84.2% on MMMU". Stacked bar (thinking 9.8 + non-thinking 74.4 = 84.2); chart labelValue 84.2.

打开官方来源

mmmu-pro 78.4 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: MMMU Pro · figure: interactive bar chart 'MMMU Pro'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote: for MMMU Pro the standard and visual task scores are averaged. Stacked bar labelValue 78.4. new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).

打开官方来源

charxiv-reasoning 81.1 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: CharXiv-Reasoning · figure: interactive bar chart 'CharXiv-Reasoning'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Stacked bar labelValue 81.1. charxiv-reasoning id already introduced by a prior batch; reused.

打开官方来源

erqa 65.7 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: ERQA · figure: interactive bar chart 'ERQA'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

erqa id already introduced by a prior batch; reused. Stacked bar labelValue 65.7.

打开官方来源

collie 99 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: COLLIE · figure: interactive bar chart 'COLLIE'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

collie id already introduced by a prior batch; reused. Stacked bar labelValue 99.

打开官方来源

video-mmmu 84.6 模型 gpt-5 · 版本 max frame 256 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: VideoMMMU · figure: interactive bar chart 'VideoMMMU'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Stacked bar labelValue 84.6. video-mmmu id introduced by the international batch; reused.

打开官方来源

healthbench 67.2 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: HealthBench · figure: interactive bar chart 'HealthBench'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: healthbench not yet in data/benchmarks/. Prose: GPT-5 scores significantly higher than any prior model on HealthBench (realistic health conversations with physician-defined rubric). Stacked bar labelValue 67.2.

打开官方来源

healthbench 46.2 模型 gpt-5 · 版本 Hard · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: HealthBench Hard · figure: interactive bar chart 'HealthBench Hard'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: "46.2% on HealthBench Hard". Stacked bar labelValue 46.2. new-benchmark: healthbench not yet in data/benchmarks/ (prior batches have not migrated it).

打开官方来源

aime-25 61.9 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "AIME 2025" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart (no tools), without thinking bar.

打开官方来源

aime-25 71 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "AIME 2025" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart (python), without thinking bar.

打开官方来源

aider 26.7 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.

打开官方来源

mmmu 74.4 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "MMMU" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.

打开官方来源

mmmu-pro 62.7 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "MMMU Pro" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar (page footnote ***).

打开官方来源

charxiv-reasoning 57.8 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "CharXiv-Reasoning" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.

打开官方来源

erqa 42 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "ERQA" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.

打开官方来源

healthbench 54.3 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "HealthBench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.

打开官方来源

video-mmmu 61.6 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "VideoMMMU" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar (max frame 256).

打开官方来源

collie 70.5 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "COLLIE" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.

打开官方来源

gpqa 77.8 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "GPQA Diamond" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart (no tools), without thinking bar.

打开官方来源

hlehle 6.3 模型 gpt-5 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "Humanity's Last Exam (Full Set)*" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart (no tools), without thinking bar.

打开官方来源

tau2-bench 0.55 模型 gpt-5 · 版本 airline, without thinking · 指标 pass_at_1 · 单位 rate 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

tau2-bench 0.076 模型 gpt-5 · 版本 airline, with thinking · 指标 pass_at_1 · 单位 rate 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

tau2-bench 0.728 模型 gpt-5 · 版本 retail, without thinking · 指标 pass_at_1 · 单位 rate 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

tau2-bench 0.083 模型 gpt-5 · 版本 retail, with thinking · 指标 pass_at_1 · 单位 rate 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

tau2-bench 0.386 模型 gpt-5 · 版本 telecom, without thinking · 指标 pass_at_1 · 单位 rate 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

tau2-bench 0.581 模型 gpt-5 · 版本 telecom, with thinking · 指标 pass_at_1 · 单位 rate 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

healthbench 1.6 模型 gpt-5 · 版本 Hard Hallucinations · 指标 inaccuracy_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - health · figure: vega chart "HealthBench Hard Hallucinations" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better (inaccuracies on challenging health conversations).

打开官方来源

healthbench 3.6 模型 gpt-5 · 版本 Hard Hallucinations · 指标 inaccuracy_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - health · figure: vega chart "HealthBench Hard Hallucinations" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better (inaccuracies on challenging health conversations).

打开官方来源

factscore 0.007 模型 gpt-5 · 版本 LongFact-Concepts · 指标 hallucination_rate · 单位 rate 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better. GPT-5 bar is the with-thinking configuration.

打开官方来源

factscore 0.008 模型 gpt-5 · 版本 LongFact-Objects · 指标 hallucination_rate · 单位 rate 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better. GPT-5 bar is the with-thinking configuration.

打开官方来源

factscore 0.01 模型 gpt-5 · 版本 FActScore · 指标 hallucination_rate · 单位 rate 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better. GPT-5 bar is the with-thinking configuration.

打开官方来源

economically-important-tasks 47.1% 模型 gpt-5 · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - economically important tasks · figure: vega chart "Economically important tasks" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: economically-important-tasks not yet in data/benchmarks. 2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart label is the stacked win-or-tie total (41.4% wins + 5.7% ties vs human experts; chart also marks an industry-expert baseline at 55). Page does not name this eval; not mapped to gdpval.

打开官方来源

GPT-5 Pro

GPT-5 Pro 是通过并行测试时计算获得更长推理的版本,发布文称其取代 o3-pro,并在 1000+ 真实推理 prompt 中获 67.8% 专家偏好(相对 GPT-5 thinking 少 22% major errors)。评测亮点:AIME 2025(python)100、GPQA Diamond 88.4(无工具)、HLE 全集 42(python + search with blocklist)。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

aime-25 96.7 模型 gpt-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

aime-25 100 模型 gpt-5-pro · 版本 with python tool · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Footnote on page: tool-assisted AIME results must not be compared with no-tools models.

打开官方来源

frontiermath 32.1 模型 gpt-5-pro · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hmmt25 100 模型 gpt-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: HMMT · figure: interactive bar chart 'HMMT'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark id usage: hmmt-25 already introduced by a prior batch; reused here. Harvard-MIT Mathematics Tournament (Feb 2025).

打开官方来源

gpqa 88.4 模型 gpt-5-pro · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: GPT-5 Pro "sets a new record on GPQA with 88.4% without tools".

打开官方来源

gpqa 89.4 模型 gpt-5-pro · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 30.7 模型 gpt-5-pro · 版本 Full Set · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

aime-25 98.4 模型 o3 · 版本 with python tool · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Cross-checks exactly with the 98.4% pass@1 printed on OpenAI's own o3/o4-mini release page.

打开官方来源

aime-25 88.9 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Cross-checks exactly with the o3/o4-mini release page chart (88.9 no tools).

打开官方来源

aime-25 42.1 模型 gpt-4o · 版本 with python tool · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 27.4 模型 chatgpt-agent · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "browser + computer + terminal",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

OpenAI's own agent product included in its release chart.

打开官方来源

frontiermath 19.3 模型 o4-mini · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

frontiermath 15.8 模型 o3 · 版本 Tier 1-3 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hmmt25 93.3 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: HMMT · figure: interactive bar chart 'HMMT'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

gpqa 83.3 模型 o3 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Cross-checks exactly with the o3/o4-mini release page chart (83.3 no tools).

打开官方来源

gpqa 70.1 模型 gpt-4o · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

hlehle 41.6 模型 chatgpt-agent · 版本 Full Set · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "browser + computer + terminal",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

OpenAI's own agent product included in its release chart.

打开官方来源

hlehle 14.7 模型 o3 · 版本 Full Set · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Differs from the 20.32 no-tools value on the o3/o4-mini release page - different eval configuration or run; do not average or silently reconcile.

打开官方来源

hlehle 5.3 模型 gpt-4o · 版本 Full Set · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

swebench 69.1 模型 o3 · 版本 Verified (fixed subset n=477) · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: SWE-bench Verified (n=477) · figure: interactive bar chart 'SWE-bench Verified (n=477)'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Cross-checks exactly with the o3/o4-mini release page (69.1, n=477 subset).

打开官方来源

swebench 30.8 模型 gpt-4o · 版本 Verified (fixed subset n=477) · 指标 resolved_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: SWE-bench Verified (n=477) · figure: interactive bar chart 'SWE-bench Verified (n=477)'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

aider 79.6 模型 o3 · 版本 Polyglot · 指标 percent_solved · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Aider Polyglot · figure: interactive bar chart 'Aider Polyglot'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

aider 25.8 模型 gpt-4o · 版本 Polyglot · 指标 percent_solved · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Aider Polyglot · figure: interactive bar chart 'Aider Polyglot'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

multichallenge 60.4 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Scale MultiChallenge · figure: interactive bar chart 'Scale MultiChallenge'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Differs from the 56.51 value on the o3/o4-mini release page; different run or protocol revision, do not reconcile silently.

打开官方来源

multichallenge 40.3 模型 gpt-4o · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: Scale MultiChallenge · figure: interactive bar chart 'Scale MultiChallenge'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

browsecomp 68.9 模型 chatgpt-agent · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: BrowseComp · figure: interactive bar chart 'BrowseComp'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

OpenAI's own agent product included in its release chart.

打开官方来源

browsecomp 49.7 模型 o3 · 版本 with python + browsing tools · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: BrowseComp · figure: interactive bar chart 'BrowseComp'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": "python + browsing",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Cross-checks exactly with the o3/o4-mini release page (49.7 with python+browsing).

打开官方来源

mmmu 82.9 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: MMMU · figure: interactive bar chart 'MMMU'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Cross-checks exactly with the o3/o4-mini release page (82.9).

打开官方来源

mmmu 72.2 模型 gpt-4o · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: MMMU · figure: interactive bar chart 'MMMU'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

mmmu-pro 76.4 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: MMMU Pro · figure: interactive bar chart 'MMMU Pro'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).

打开官方来源

charxiv-reasoning 78.6 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: CharXiv-Reasoning · figure: interactive bar chart 'CharXiv-Reasoning'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Cross-checks exactly with the o3/o4-mini release page (78.6).

打开官方来源

erqa 64 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: ERQA · figure: interactive bar chart 'ERQA'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

collie 98.4 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: COLLIE · figure: interactive bar chart 'COLLIE'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

healthbench 59.8 模型 o3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: HealthBench · figure: interactive bar chart 'HealthBench'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: healthbench not yet in data/benchmarks/.

打开官方来源

healthbench 31.6 模型 o3 · 版本 Hard · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark charts (math / science / coding / multimodal sections) · row: HealthBench Hard · figure: interactive bar chart 'HealthBench Hard'; values machine-read from page RSC vega-lite payload

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: healthbench not yet in data/benchmarks/ (prior batches have not migrated it).

打开官方来源

video-mmmu 83.3 模型 o3 · 版本 max frame 256 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "VideoMMMU" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 max frame 256.

打开官方来源

video-mmmu 61.2 模型 gpt-4o · 版本 max frame 256 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "VideoMMMU" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 max frame 256.

打开官方来源

mmmu-pro 59.9 模型 gpt-4o · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "MMMU Pro" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。

打开官方来源

charxiv-reasoning 58.8 模型 gpt-4o · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "CharXiv-Reasoning" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。

打开官方来源

erqa 35.2 模型 gpt-4o · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "ERQA" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。

打开官方来源

healthbench 32 模型 gpt-4o · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "HealthBench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。

打开官方来源

healthbench 0 模型 gpt-4o · 版本 Hard · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "HealthBench Hard" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 chart prints 0 for GPT-4o.

打开官方来源

hlehle 24.3 模型 o3 · 版本 Full Set, python + browser · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "Humanity's Last Exam (Full Set)*" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。

打开官方来源

hlehle 23 模型 chatgpt-agent · 版本 Full Set (no tools) · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts · figure: vega chart "Humanity's Last Exam (Full Set)*" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。

打开官方来源

tau2-bench 0.648 模型 o3 · 版本 airline · 指标 pass_at_1 · 单位 rate 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

tau2-bench 0.802 模型 o3 · 版本 retail · 指标 pass_at_1 · 单位 rate 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

tau2-bench 0.582 模型 o3 · 版本 telecom · 指标 pass_at_1 · 单位 rate 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

tau2-bench 0.455 模型 gpt-4o · 版本 airline · 指标 pass_at_1 · 单位 rate 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

tau2-bench 0.634 模型 gpt-4o · 版本 retail · 指标 pass_at_1 · 单位 rate 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

tau2-bench 0.235 模型 gpt-4o · 版本 telecom · 指标 pass_at_1 · 单位 rate 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.

打开官方来源

healthbench 12.9 模型 o3 · 版本 Hard Hallucinations · 指标 inaccuracy_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - health · figure: vega chart "HealthBench Hard Hallucinations" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better (inaccuracies on challenging health conversations).

打开官方来源

healthbench 15.8 模型 gpt-4o · 版本 Hard Hallucinations · 指标 inaccuracy_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - health · figure: vega chart "HealthBench Hard Hallucinations" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better (inaccuracies on challenging health conversations).

打开官方来源

factscore 0.045 模型 o3 · 版本 LongFact-Concepts · 指标 hallucination_rate · 单位 rate 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better.

打开官方来源

factscore 0.051 模型 o3 · 版本 LongFact-Objects · 指标 hallucination_rate · 单位 rate 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better.

打开官方来源

factscore 0.057 模型 o3 · 版本 FActScore · 指标 hallucination_rate · 单位 rate 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better.

打开官方来源

economically-important-tasks 33.5% 模型 o3 · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - economically important tasks · figure: vega chart "Economically important tasks" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: economically-important-tasks not yet in data/benchmarks. 2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart label is the stacked win-or-tie total (26.9% wins + 6.6% ties vs human experts; chart also marks an industry-expert baseline at 55). Page does not name this eval; not mapped to gdpval.

打开官方来源

economically-important-tasks 43.5% 模型 chatgpt-agent · 版本 未说明 · 指标 win_or_tie_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation charts - economically important tasks · figure: vega chart "Economically important tasks" (dataset values machine-read from page payload, 2026-09-01)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: economically-important-tasks not yet in data/benchmarks. 2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart label is the stacked win-or-tie total (36.0% wins + 7.5% ties vs human experts; chart also marks an industry-expert baseline at 55). Page does not name this eval; not mapped to gdpval.

打开官方来源