← 模型目录

Claude Sonnet 5

Anthropic · 2026-06-30 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Claude Sonnet 5

发布文将其定位为最具智能体能力(most agentic)的 Sonnet,支持可选努力档位、默认开启网络安全防护,并作为 Free/Pro 默认模型全量提供。评测集中于智能体编码、计算机使用、知识工作与安全评估,亮点为 Terminal-Bench 2.1 80.4% 与 OSWorld-Verified 81.2%。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 2 / 输出 10 · 初上市定价 USD 2/10,原计划 2026-09-01 起的 USD 3/15 未生效,现已转为永久定价

本变体的评测证据

firefox-147-exploit 0.0% 模型 claude-sonnet-5 · 版本 working exploit rate · 指标 working_exploit_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Safety evaluations section · row: Firefox 147 exploit development (Mozilla collaboration) · quote_snippet: Sonnet 5 was never able to develop a full working exploit [...] both scored 0.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "evaluation developed in collaboration with Mozilla; all vulnerabilities patched in Firefox 148; models tested without safeguards"
}

new-benchmark: firefox-147-exploit not yet in data/benchmarks/. Left bar = working exploit, right bar = partial success (chart image ee9944c8). Sonnet 5 shows a slightly higher partial-success rate than Sonnet 4.6, attributed by Anthropic to general intelligence rather than specific training.

打开官方来源

swebench-pro 63.2 模型 claude-sonnet-5 · 版本 Public · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Agentic coding - SWE-bench Pro · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 63.2 vs Sonnet 4.6 58.1, Opus 4.8 69.2 (cross-checks exactly with the Opus 4.8 page table).

打开官方来源

terminalbench 80.4 模型 claude-sonnet-5 · 版本 2.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Agentic coding - Terminal-Bench 2.1 · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 80.4 vs Sonnet 4.6 67.0, Opus 4.8 82.7.

打开官方来源

hlehle 43.2 模型 claude-sonnet-5 · 版本 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Multidisciplinary reasoning - Humanity's Last Exam (no tools) · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 43.2 vs Sonnet 4.6 34.6, Opus 4.8 49.8 (Opus cross-checks exactly with its own page).

打开官方来源

hlehle 57.4 模型 claude-sonnet-5 · 版本 with tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Multidisciplinary reasoning - Humanity's Last Exam (with tools) · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 57.4 vs Sonnet 4.6 46.8, Opus 4.8 57.9.

打开官方来源

osworld 81.2 模型 claude-sonnet-5 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Computer use - OSWorld-Verified · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 81.2 vs Sonnet 4.6 78.5, Opus 4.8 83.4.

打开官方来源

gdpval 1618 模型 claude-sonnet-5 · 版本 AA v2 (Elo) · 指标 未说明 · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Knowledge work - GDPval-AA v2 · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 1618 vs Sonnet 4.6 1395, Opus 4.8 1615. The v2 suffix appears only on this page (Opus 4.8 page shows GDPval-AA without version label, Sonnet 5 at 1615 vs Opus 4.8 page 1890 for Opus) - version differences flagged, do not merge silently.

打开官方来源

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

firefox-147-exploit 0.0% 模型 claude-sonnet-4-6 · 版本 working exploit rate · 指标 working_exploit_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Safety evaluations section · row: Firefox 147 exploit development · quote_snippet: Neither of the Sonnet models could successfully develop a working exploit (both scored 0.0%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: firefox-147-exploit not yet in data/benchmarks/.

打开官方来源