GPT-5 / GPT-5 Pro
OpenAI · 2025-08-07 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GPT-5
GPT-5 是 OpenAI 的统一思考模型,发布文称其在 AIME 2025(无工具 94.6%)、SWE-bench Verified 74.9%、HealthBench Hard 46.2% 等刷新 SOTA,并大幅降低事实错误与谄媚率。评测覆盖竞赛数学、编码、指令跟随与代理、多模态与健康:GPQA Diamond 88.4%、MMMU 84.2%、BrowseComp 54.9%。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose anchor: "94.6% on AIME 2025 without tools". Stacked bar (thinking 32.7 + non-thinking 61.9 = 94.6); chart labelValue 94.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote on page: tool-assisted AIME results must not be compared with no-tools models.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: HMMT · figure: interactive bar chart 'HMMT'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: HMMT · figure: interactive bar chart 'HMMT'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Stacked bar (thinking 7.9 + non-thinking 77.8 = 85.7); chart labelValue 85.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Stacked bar (thinking 18.5 + non-thinking 6.3 = 24.8); chart labelValue 24.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python + search with blocklist",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: SWE-bench Verified (n=477) · figure: interactive bar chart 'SWE-bench Verified (n=477)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high (default); effort line chart: low 69.1 / medium 72.4 / high 74.9",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: "74.9% on SWE-bench Verified". Companion line chart (tokens vs accuracy) gives GPT-5 low/medium/high 69.1/72.4/74.9 and o3 low/medium/high 63.8/67.1/69.1 at 4202-13741 output tokens. Footnote: fixed subset of n=477 verified tasks validated on OpenAI internal infrastructure.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: SWE-bench Verified (n=477) · figure: interactive bar chart 'SWE-bench Verified (n=477)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Aider Polyglot · figure: interactive bar chart 'Aider Polyglot'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: "88% on Aider-Polyglot". Stacked bar (thinking 61.3 + non-thinking 26.7 = 88); chart labelValue 88.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Scale MultiChallenge · figure: interactive bar chart 'Scale MultiChallenge'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Stacked bar (thinking 14.7 + non-thinking 54.9 = 69.6); chart labelValue 69.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: BrowseComp · figure: interactive bar chart 'BrowseComp'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: MMMU · figure: interactive bar chart 'MMMU'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: "84.2% on MMMU". Stacked bar (thinking 9.8 + non-thinking 74.4 = 84.2); chart labelValue 84.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: MMMU Pro · figure: interactive bar chart 'MMMU Pro'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote: for MMMU Pro the standard and visual task scores are averaged. Stacked bar labelValue 78.4. new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: CharXiv-Reasoning · figure: interactive bar chart 'CharXiv-Reasoning'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Stacked bar labelValue 81.1. charxiv-reasoning id already introduced by a prior batch; reused.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: ERQA · figure: interactive bar chart 'ERQA'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}erqa id already introduced by a prior batch; reused. Stacked bar labelValue 65.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: COLLIE · figure: interactive bar chart 'COLLIE'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}collie id already introduced by a prior batch; reused. Stacked bar labelValue 99.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: VideoMMMU · figure: interactive bar chart 'VideoMMMU'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Stacked bar labelValue 84.6. video-mmmu id introduced by the international batch; reused.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: HealthBench · figure: interactive bar chart 'HealthBench'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: healthbench not yet in data/benchmarks/. Prose: GPT-5 scores significantly higher than any prior model on HealthBench (realistic health conversations with physician-defined rubric). Stacked bar labelValue 67.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: HealthBench Hard · figure: interactive bar chart 'HealthBench Hard'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: "46.2% on HealthBench Hard". Stacked bar labelValue 46.2. new-benchmark: healthbench not yet in data/benchmarks/ (prior batches have not migrated it).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "AIME 2025" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart (no tools), without thinking bar.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "AIME 2025" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart (python), without thinking bar.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "Aider Polyglot" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "MMMU" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "MMMU Pro" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar (page footnote ***).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "CharXiv-Reasoning" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "ERQA" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "HealthBench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "VideoMMMU" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar (max frame 256).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "COLLIE" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart without thinking bar.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "GPQA Diamond" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart (no tools), without thinking bar.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "Humanity's Last Exam (Full Set)*" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart (no tools), without thinking bar.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - health · figure: vega chart "HealthBench Hard Hallucinations" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better (inaccuracies on challenging health conversations).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - health · figure: vega chart "HealthBench Hard Hallucinations" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better (inaccuracies on challenging health conversations).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better. GPT-5 bar is the with-thinking configuration.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better. GPT-5 bar is the with-thinking configuration.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better. GPT-5 bar is the with-thinking configuration.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - economically important tasks · figure: vega chart "Economically important tasks" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: economically-important-tasks not yet in data/benchmarks. 2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart label is the stacked win-or-tie total (41.4% wins + 5.7% ties vs human experts; chart also marks an industry-expert baseline at 55). Page does not name this eval; not mapped to gdpval.
GPT-5 Pro
GPT-5 Pro 是通过并行测试时计算获得更长推理的版本,发布文称其取代 o3-pro,并在 1000+ 真实推理 prompt 中获 67.8% 专家偏好(相对 GPT-5 thinking 少 22% major errors)。评测亮点:AIME 2025(python)100、GPQA Diamond 88.4(无工具)、HLE 全集 42(python + search with blocklist)。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Footnote on page: tool-assisted AIME results must not be compared with no-tools models.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: HMMT · figure: interactive bar chart 'HMMT'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark id usage: hmmt-25 already introduced by a prior batch; reused here. Harvard-MIT Mathematics Tournament (Feb 2025).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: GPT-5 Pro "sets a new record on GPQA with 88.4% without tools".
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python + search with blocklist",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Cross-checks exactly with the 98.4% pass@1 printed on OpenAI's own o3/o4-mini release page.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Cross-checks exactly with the o3/o4-mini release page chart (88.9 no tools).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: AIME 2025 · figure: interactive bar chart 'AIME 2025'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "browser + computer + terminal",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}OpenAI's own agent product included in its release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: FrontierMath, Tier 1-3 · figure: interactive bar chart 'FrontierMath, Tier 1-3'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: HMMT · figure: interactive bar chart 'HMMT'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Cross-checks exactly with the o3/o4-mini release page chart (83.3 no tools).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: GPQA Diamond · figure: interactive bar chart 'GPQA Diamond'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "browser + computer + terminal",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}OpenAI's own agent product included in its release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Differs from the 20.32 no-tools value on the o3/o4-mini release page - different eval configuration or run; do not average or silently reconcile.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Humanity's Last Exam (Full Set) · figure: interactive bar chart 'Humanity's Last Exam (Full Set)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: SWE-bench Verified (n=477) · figure: interactive bar chart 'SWE-bench Verified (n=477)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Cross-checks exactly with the o3/o4-mini release page (69.1, n=477 subset).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: SWE-bench Verified (n=477) · figure: interactive bar chart 'SWE-bench Verified (n=477)'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Aider Polyglot · figure: interactive bar chart 'Aider Polyglot'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Aider Polyglot · figure: interactive bar chart 'Aider Polyglot'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Scale MultiChallenge · figure: interactive bar chart 'Scale MultiChallenge'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Differs from the 56.51 value on the o3/o4-mini release page; different run or protocol revision, do not reconcile silently.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: Scale MultiChallenge · figure: interactive bar chart 'Scale MultiChallenge'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: BrowseComp · figure: interactive bar chart 'BrowseComp'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}OpenAI's own agent product included in its release chart.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: BrowseComp · figure: interactive bar chart 'BrowseComp'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": "python + browsing",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Cross-checks exactly with the o3/o4-mini release page (49.7 with python+browsing).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: MMMU · figure: interactive bar chart 'MMMU'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Cross-checks exactly with the o3/o4-mini release page (82.9).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: MMMU · figure: interactive bar chart 'MMMU'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: MMMU Pro · figure: interactive bar chart 'MMMU Pro'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: CharXiv-Reasoning · figure: interactive bar chart 'CharXiv-Reasoning'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Cross-checks exactly with the o3/o4-mini release page (78.6).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: ERQA · figure: interactive bar chart 'ERQA'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: COLLIE · figure: interactive bar chart 'COLLIE'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: HealthBench · figure: interactive bar chart 'HealthBench'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: healthbench not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark charts (math / science / coding / multimodal sections) · row: HealthBench Hard · figure: interactive bar chart 'HealthBench Hard'; values machine-read from page RSC vega-lite payload
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: healthbench not yet in data/benchmarks/ (prior batches have not migrated it).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "VideoMMMU" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 max frame 256.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "VideoMMMU" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 max frame 256.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "MMMU Pro" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "CharXiv-Reasoning" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "ERQA" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "HealthBench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "HealthBench Hard" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 chart prints 0 for GPT-4o.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "Humanity's Last Exam (Full Set)*" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts · figure: vega chart "Humanity's Last Exam (Full Set)*" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - function calling · figure: vega chart "Tau2-bench" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart prints rates as fractions (0-1); per-domain bars. GPT-5 bars are split by thinking mode.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - health · figure: vega chart "HealthBench Hard Hallucinations" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better (inaccuracies on challenging health conversations).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - health · figure: vega chart "HealthBench Hard Hallucinations" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better (inaccuracies on challenging health conversations).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - factuality · figure: vega chart "Hallucination rate on open-source prompts" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Lower is better.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - economically important tasks · figure: vega chart "Economically important tasks" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: economically-important-tasks not yet in data/benchmarks. 2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart label is the stacked win-or-tie total (26.9% wins + 6.6% ties vs human experts; chart also marks an industry-expert baseline at 55). Page does not name this eval; not mapped to gdpval.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation charts - economically important tasks · figure: vega chart "Economically important tasks" (dataset values machine-read from page payload, 2026-09-01)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: economically-important-tasks not yet in data/benchmarks. 2026-09-01 审计补录:图内数据点取自页面 vega 图表数据。 Chart label is the stacked win-or-tie total (36.0% wins + 7.5% ties vs human experts; chart also marks an industry-expert baseline at 55). Page does not name this eval; not mapped to gdpval.