OpenAI o1
OpenAI · 2024-09-12 · 类别未确认
OpenAI o1
OpenAI 首款强化学习驱动思维链(Chain of Thought)推理模型,在科学、编程与数学竞争性基准测试中展现出突破性的深度思考能力。
- 输入模态
- 文本 / 图像
- 上下文
- 128K tokens
- 参数
- 未公开
- 价格(每百万 tokens)
- 输入 15 / 输出 60
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。
未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals section - evaluation table · table: SSR evaluation table, columns: gpt-4o / o1-preview / o1 · row: MMLU - pass@1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计重写:自 adoption[] 迁移的空壳行按页面评测表填值(o1-preview 列 pass@1)。同行 gpt-4o 88.0 / o1 90.8 未入行(legacy 基线不得扩容,见 release notes)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals section - evaluation table · table: SSR evaluation table, columns: gpt-4o / o1-preview / o1 · row: GPQA Diamond - pass@1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计重写:页面评测表 GPQA Diamond pass@1(o1-preview 列);cons@64 78.3 记于此备注。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals section - evaluation table · table: SSR evaluation table, columns: gpt-4o / o1-preview / o1 · row: Competition Math AIME (2024) - pass@1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计重写:页面评测表 AIME (2024) pass@1(o1-preview 列);cons@64 56.7 记于此备注。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals section - evaluation table · table: SSR evaluation table, columns: gpt-4o / o1-preview / o1 · row: MATH - pass@1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 审计重写:原迁移行记为 math500,但页面行标签仅为 MATH 且未注明子集,故 benchmark_id 改为 math。o1-preview 列 pass@1。