GPT-6 Sol / GPT-6 Luna
OpenAI · 2026-09-22 · 通用模型
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GPT-6 Sol
OpenAI 将 GPT-6 Sol 定位为日常复杂高价值任务的旗舰级高性价比工作马,以 GPT-5.6 Sol 一半的价格提供接近 Astra 的智能;收录评测覆盖企业自动化、软件工程、计算机操作与专业考试,AutomationBench 达 33.2%、DeepSWE 达 68.8%。
- 输入模态
- 文本 / 图像 / 代码 / 计算机操作
- 上下文
- 1.05M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 2 / 输出 10 · API 价格相比 GPT-5.6 Sol 降价 50%;缓存命中读取享 90% 折扣
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional work · table: Model (and effort) · row: GPT-6 Sol (xhigh) · figure: AutomationBench chart (cost: $0.27; score: 33%) · quote_snippet: On AutomationBench, a test of business workflows across apps, GPT‑6 Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of Opus 5’s cost per task.
{
"harness": "AutomationBench 1.0.6",
"tools": [
"browser",
"bash",
"python"
],
"shots": null,
"reasoning_effort": "xhigh",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方博文与图表数据:在 AutomationBench 1.0.6(涵盖销售、营销、运营、财务等 47 项工具)取得 33.2%,单任务成本 $0.27。超越 Claude Opus 5 max effort (26.9%, $3.05) 与 Claude Fable 5.1 w/ fallback (31.4%, $2.45)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional work · figure: Agents' Last Exam V1 chart (cost: $2.93; score: 56%) · quote_snippet: On Agents’ Last Exam, which evaluates agents on complex professional workflows, GPT‑6 Sol at max effort scores 56.4%, above Claude Opus 5’s highest score in the evaluation at 60% lower cost per task.
{
"harness": "Agents’ Last Exam V1",
"tools": [
"browser",
"bash",
"python"
],
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方博文与图表数据:跨 55 个行业专业工作流任务获得 56.4%,成本 $2.93,超越 Claude Opus 5 最高成绩 (55.9%, $7.29) 且成本降低 60%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding · figure: DeepSWE 1.1 chart (cost: $2.74; score: 69%) · quote_snippet: On DeepSWE v1.1, which tests performance on complex software-engineering tasks in real codebases, GPT‑6 Sol at max effort scores 68.8%, within 1.1 percentage points of Claude Fable 5’s highest score in the evaluation—69.9% at xhigh effort—at approximately 80% lower cost per task.
{
"harness": "DeepSWE 1.1",
"tools": [
"bash",
"code_editor"
],
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方博文与图表数据:在 DeepSWE 1.1 软件工程长任务中获得 68.8%,成本 $2.74,距离 Claude Fable 5 最高分 (69.9% at xhigh, $13.41) 仅 1.1 个百分点,单任务成本降低约 80%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding · figure: FrontierCode 1.1 Main chart (cost: $2.14; score: 49%) · quote_snippet: On FrontierCode, which evaluates whether coding agents produce changes ready to merge into real codebases, GPT‑6 Sol improves substantially over GPT‑5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.
{
"harness": "FrontierCode 1.1 Main",
"tools": [
"bash",
"code_editor",
"git"
],
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方博文与图表数据:评测代码产出的可合并性(mergeability)。GPT-6 Sol max effort 取得 49.3%(成本 $2.14),xhigh 取得 48.4%(成本 $1.37),以远低于 Claude Fable 5.1 xhigh ($9.27, 48.7%) 的成本匹敌其表现。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use · figure: OSWorld 2.0 offline chart (cost: $2.21; score: 61%) · quote_snippet: On OSWorld 2.0 offline, GPT‑6 Sol at xhigh effort achieves a similar score to Claude Opus 5 at medium effort—60.5% versus 60.3%—at approximately 80% lower cost per task.
{
"harness": "OSWorld 2.0 offline (v2026.08.08 partial reward)",
"tools": [
"gui_screen",
"mouse",
"keyboard"
],
"shots": null,
"reasoning_effort": "xhigh",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方博文与图表数据:评测长周期真实计算机操控任务。GPT-6 Sol xhigh effort 取得 60.5%(单任务成本 $2.21),匹敌 Claude Opus 5 medium effort (60.3%, $12.67),成本降低约 80%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use · figure: OSWorld 2.0 offline chart (cost: $3.25; score: 64%)
{
"harness": "OSWorld 2.0 offline (v2026.08.08 partial reward)",
"tools": [
"gui_screen",
"mouse",
"keyboard"
],
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方图表数据:GPT-6 Sol 在 max effort 达到 64.4%,单任务成本 $3.25。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Factuality · figure: Answers with any factual error chart (cost: $0.13; score: 5%) · quote_snippet: On our internal factuality evaluation, which is based on de-identified real-world conversations where users flagged mistakes by our models, GPT‑6 Sol makes about half as many mistakes as its predecessor, approaching Astra-level reliability at much lower cost.
{
"harness": "OpenAI internal factuality evaluation on de-identified error-flagged ChatGPT conversations",
"tools": null,
"shots": null,
"reasoning_effort": "xhigh",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: OpenAI 官方内部事实准确度评测。衡量含有事实错误的回答比例(Answers with any factual error,越低越好)。GPT-6 Sol xhigh effort 为 4.5%,max effort 为 4.6%,较 GPT-5.6 Sol (8.4%) 错误率减半,接近 Astra 水平 (4.0%)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Continuing to improve alignment · figure: Coding deception chart (score: 1.3%) · quote_snippet: In our internal coding deception evaluation, AI agents are given tasks deliberately selected to elicit dishonesty. In typical usage, deception is much rarer. Deception rate measures the fraction of answers with any detected deception. Effort was set to maximum.
{
"harness": "OpenAI internal coding deception evaluation",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: OpenAI 官方对齐评测(内部编码欺瞒率)。在刻意诱导不诚实行为的代码任务中检测欺瞒回答比例(Deception rate,越低越好)。GPT-6 Sol max effort 为 1.3%,显著低于 GPT-5.6 Sol 的 10.4%。
GPT-6 Luna
OpenAI 将 GPT-6 Luna 定位为极致高吞吐、超低成本的轻量智能模型,面向高频摘要、数据提取、分类与路由任务;收录评测显示其在高推理预算下可匹敌前代 Sol,DeepSWE 达 66.6%、OSWorld 达 52.7%。
- 输入模态
- 文本 / 图像 / 代码 / 计算机操作
- 上下文
- 1.05M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 0.1 / 输出 0.5 · API 价格相比 GPT-5.6 Luna 降价 50%;缓存命中读取享 90% 折扣
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional work · figure: AutomationBench chart (cost: $0.021; score: 14%) · quote_snippet: At high effort, GPT‑6 Luna improves on its predecessor by 5.4 percentage points at 58% lower cost per task.
{
"harness": "AutomationBench 1.0.6",
"tools": [
"browser",
"bash",
"python"
],
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方博文与图表数据:在 high effort 获得 14.5%,成本仅 $0.021,相比前代 GPT-5.6 Luna (9.1%) 提升 5.4 个百分点且成本降低 58%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional work · figure: AutomationBench chart (cost: $0.037; score: 21%)
{
"harness": "AutomationBench 1.0.6",
"tools": [
"browser",
"bash",
"python"
],
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方图表数据:GPT-6 Luna 在 max effort 下取得 20.7%,单任务成本仅 $0.037。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional work · figure: Agents' Last Exam V1 chart (cost: $0.15; score: 51%)
{
"harness": "Agents’ Last Exam V1",
"tools": [
"browser",
"bash",
"python"
],
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方图表数据:GPT-6 Luna 在 max effort 下取得 50.9%,成本仅 $0.15,接近前代旗舰 GPT-5.6 Sol 的最高得分 (53.6%)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding · figure: DeepSWE 1.1 chart (cost: $0.22; score: 67%) · quote_snippet: GPT‑6 Luna at max effort scores 66.6%, comparable to Claude Opus 5 and Fable 5 at medium effort. In these comparisons, Luna costs 93% less per task than Opus 5 and 96% less than Fable 5.
{
"harness": "DeepSWE 1.1",
"tools": [
"bash",
"code_editor"
],
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方博文与图表数据:GPT-6 Luna 在 max effort 取得 66.6%,单任务成本仅 $0.22,匹敌 Claude Opus 5 与 Fable 5 在 medium effort 的表现 (68.9% / 65.4%),成本降低 93%~96%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding · figure: FrontierCode 1.1 Main chart (cost: $0.11; score: 42%)
{
"harness": "FrontierCode 1.1 Main",
"tools": [
"bash",
"code_editor",
"git"
],
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方图表数据:GPT-6 Luna 在 max effort 下取得 42.4%,单任务成本仅 $0.11。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Computer use · figure: OSWorld 2.0 offline chart (cost: $0.27; score: 53%) · quote_snippet: GPT‑6 Luna (max) is able to exceed GPT‑5.6 Sol (medium) at one tenth of its cost.
{
"harness": "OSWorld 2.0 offline (v2026.08.08 partial reward)",
"tools": [
"gui_screen",
"mouse",
"keyboard"
],
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}官方博文与图表数据:GPT-6 Luna (max effort) 取得 52.7%,成本 $0.27,超越前代主力 GPT-5.6 Sol medium effort (49.7%, $2.73) 且成本仅为其十分之一。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Factuality · figure: Answers with any factual error chart (cost: $0.012; score: 8%) · quote_snippet: GPT‑6 Luna also improves substantially; at higher effort levels it matches GPT‑5.6 Sol at about a hundredth its cost.
{
"harness": "OpenAI internal factuality evaluation on de-identified error-flagged ChatGPT conversations",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: OpenAI 官方内部事实准确度评测。GPT-6 Luna 在 max effort 下事实错误率为 7.6%(成本 $0.012),优于 GPT-5.6 Sol 的 max effort 8.5%(成本 $0.87),以百分之一的成本达到前代 Sol 水平。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Continuing to improve alignment · figure: Coding deception chart (score: 2.8%) · quote_snippet: In our internal coding deception evaluation, AI agents are given tasks deliberately selected to elicit dishonesty. In typical usage, deception is much rarer. Deception rate measures the fraction of answers with any detected deception. Effort was set to maximum.
{
"harness": "OpenAI internal coding deception evaluation",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: OpenAI 官方对齐评测(内部编码欺瞒率)。GPT-6 Luna max effort 为 2.8%,显著低于 GPT-5.6 Luna 的 9.5%。