← 模型目录

GPT-6.1 Sol

OpenAI · 2026-09-29 · 通用模型

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GPT-6.1 Sol

GPT-6 Sol 的升级版,定位是在智能体编程、计算机使用和专业工作上更接近 GPT-6 Astra,而标准价格仅为 Astra 的五分之一。收录的 7 项评测中 5 项为图表行(数值未读出),正文明示的事实错误率为 4.1%。

输入模态
官方资料未说明
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 2 / 输出 10 · 缓存输入 $0.10/百万 token

本变体的评测证据

deepswe 图表尚无可读数值 模型 gpt-6-1-sol · 版本 DeepSWE v1.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 各项任务上能力更强的 Sol · row: DeepSWE v1.1 · figure: 得分对单任务成本图(DeepSWE)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

图中数值未读出。正文称:GPT-6.1 Sol 以约五分之一成本达到与 GPT-6 Astra 相当水平,以更低推理强度和成本,比 GPT-6 Sol 最高得分高 6.4 个百分点。

打开官方来源

gdp-pdf 图表尚无可读数值 模型 gpt-6-1-sol · 版本 GDP.pdf · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 各项任务上能力更强的 Sol · row: GDP.pdf · figure: 得分对单任务成本图(GDP.pdf)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: GDP.pdf 衡量模型根据复杂 PDF 文档回答金融、医疗、法律等专业问题的准确度。图中数值未读出。正文称:各推理档得分均高于带回退机制的 Opus 5.5,单任务成本不到其一半;以约五分之一成本接近 GPT-6 Astra。

打开官方来源

automationbench 图表尚无可读数值 模型 gpt-6-1-sol · 版本 AutomationBench 1.0.6, medium reasoning · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 各项任务上能力更强的 Sol · row: AutomationBench 1.0.6, medium reasoning · figure: 得分对单任务成本图(AutomationBench)

{
  "harness": null,
  "tools": [
    "47 tools"
  ],
  "shots": null,
  "reasoning_effort": "medium",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

图中数值未读出。正文称:中等推理强度下比 Opus 5.5 高 2.2 个百分点、成本约三分之一,比同设置 GPT-6 Sol 高 4.8 个百分点。Claude Fable 5.1 数据点未计入回退成本(约 40% 任务发生回退)。

打开官方来源

osworld 图表尚无可读数值 模型 gpt-6-1-sol · 版本 OSWorld 2.0 offline set (v2026.08.08), partial reward, max reasoning · 指标 partial reward · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 各项任务上能力更强的 Sol · row: OSWorld 2.0 offline set (v2026.08.08), partial reward, max reasoning · figure: 得分对单任务成本图(OSWorld)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

图中数值未读出。正文称:最高推理强度下比 GPT-6 Sol 高 7 个百分点、成本不到其一半;与 Astra 差距缩小至 2.1 个百分点以内,单任务成本约为其七分之一。

打开官方来源

terminal-bench-science 图表尚无可读数值 模型 gpt-6-1-sol · 版本 Terminal-Bench Science 0.1, max reasoning · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 各项任务上能力更强的 Sol · row: Terminal-Bench Science 0.1, max reasoning · figure: 得分对单任务成本图(Terminal-Bench Science)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

图中数值未读出。正文称:最高推理强度下得分超过 GPT-6 Sol 的两倍,单任务成本不到其一半;GPT-6.1 Sol 平均单任务成本 $5.47,Opus 5.5 为 $23.21,Astra 为 $23.80;GPT-6 Astra 以 68.1% 在受测模型中排名第一(同族模型数值,不是本模型得分)。

打开官方来源

chatgpt-factual-error-rate 4.1% 模型 gpt-6-1-sol · 版本 Factual error rate on hard de-identified ChatGPT conversations, xhigh reasoning · 指标 lower is better · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 事实准确性 · row: Factual error rate on hard de-identified ChatGPT conversations, xhigh reasoning · quote_snippet: 事实错误率为 4.1%,低于 GPT‑6 Sol 的 4.5%,更接近相同设置下 Astra 的 4.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "xhigh",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: 用去标识化的 ChatGPT 对话(用户曾指出早期模型的事实错误)衡量回答含至少一处事实错误的比例,数值越低越好;提示为刻意挑选的高难度提示,不代表典型使用。GPT-6 Sol 4.5%,GPT-6 Astra 4.0%。

打开官方来源

search-tool-failure-disclosure 2.1% 模型 gpt-6-1-sol · 版本 Search tool failure non-disclosure rate, max reasoning · 指标 lower is better · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 安全部署 GPT-6.1 Sol · row: Search tool failure non-disclosure rate, max reasoning · quote_snippet: GPT‑6.1 Sol 在 2.1% 的测试案例中未披露这一问题

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: 测试智能体在搜索工具故障时是否如实告知用户,而不是给出自己认为最可能的答案,比例越低越好;测试任务刻意诱发失误,不代表典型使用。GPT-6 Sol 4.9%,GPT-6 Astra 1.5%,GPT-6 Luna 28.7%。

打开官方来源