← 模型目录

OpenAI o1

OpenAI · 2024-09-12 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口

OpenAI o1

OpenAI 首款强化学习驱动思维链(Chain of Thought)推理模型,在科学、编程与数学竞争性基准测试中展现出突破性的深度思考能力。

输入模态
文本 / 图像
上下文
128K tokens
参数
未公开
价格(每百万 tokens)
输入 15 / 输出 60

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

mmlu 92.3 模型 o1-preview · 版本 未说明 · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals section - evaluation table · table: SSR evaluation table, columns: gpt-4o / o1-preview / o1 · row: MMLU - pass@1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计重写:自 adoption[] 迁移的空壳行按页面评测表填值(o1-preview 列 pass@1)。同行 gpt-4o 88.0 / o1 90.8 未入行(legacy 基线不得扩容,见 release notes)。

打开官方来源

gpqa 73.3 模型 o1-preview · 版本 Diamond · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals section - evaluation table · table: SSR evaluation table, columns: gpt-4o / o1-preview / o1 · row: GPQA Diamond - pass@1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计重写:页面评测表 GPQA Diamond pass@1(o1-preview 列);cons@64 78.3 记于此备注。

打开官方来源

aime24 44.6 模型 o1-preview · 版本 未说明 · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals section - evaluation table · table: SSR evaluation table, columns: gpt-4o / o1-preview / o1 · row: Competition Math AIME (2024) - pass@1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计重写:页面评测表 AIME (2024) pass@1(o1-preview 列);cons@64 56.7 记于此备注。

打开官方来源

math 85.5 模型 o1-preview · 版本 未说明 · 指标 pass_at_1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals section - evaluation table · table: SSR evaluation table, columns: gpt-4o / o1-preview / o1 · row: MATH - pass@1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 审计重写:原迁移行记为 math500,但页面行标签仅为 MATH 且未注明子集,故 benchmark_id 改为 math。o1-preview 列 pass@1。

打开官方来源