Claude Sonnet 5
Anthropic · 2026-06-30 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Claude Sonnet 5
发布文将其定位为最具智能体能力(most agentic)的 Sonnet,支持可选努力档位、默认开启网络安全防护,并作为 Free/Pro 默认模型全量提供。评测集中于智能体编码、计算机使用、知识工作与安全评估,亮点为 Terminal-Bench 2.1 80.4% 与 OSWorld-Verified 81.2%。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 2 / 输出 10 · 初上市定价 USD 2/10,原计划 2026-09-01 起的 USD 3/15 未生效,现已转为永久定价
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Safety evaluations section · row: Firefox 147 exploit development (Mozilla collaboration) · quote_snippet: Sonnet 5 was never able to develop a full working exploit [...] both scored 0.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "evaluation developed in collaboration with Mozilla; all vulnerabilities patched in Firefox 148; models tested without safeguards"
}new-benchmark: firefox-147-exploit not yet in data/benchmarks/. Left bar = working exploit, right bar = partial success (chart image ee9944c8). Sonnet 5 shows a slightly higher partial-success rate than Sonnet 4.6, attributed by Anthropic to general intelligence rather than specific training.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Agentic coding - SWE-bench Pro · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 63.2 vs Sonnet 4.6 58.1, Opus 4.8 69.2 (cross-checks exactly with the Opus 4.8 page table).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Agentic coding - Terminal-Bench 2.1 · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 80.4 vs Sonnet 4.6 67.0, Opus 4.8 82.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Multidisciplinary reasoning - Humanity's Last Exam (no tools) · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 43.2 vs Sonnet 4.6 34.6, Opus 4.8 49.8 (Opus cross-checks exactly with its own page).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Multidisciplinary reasoning - Humanity's Last Exam (with tools) · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 57.4 vs Sonnet 4.6 46.8, Opus 4.8 57.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Computer use - OSWorld-Verified · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 81.2 vs Sonnet 4.6 78.5, Opus 4.8 83.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Knowledge work - GDPval-AA v2 · figure: images/02.webp(归档 benchmark 总表 2600x1234;列: Sonnet 5 | Sonnet 4.6 | Opus 4.8 (For reference)) · quote_snippet: Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 5 列,与先前 vision-read 逐格一致)。Sonnet 5 1618 vs Sonnet 4.6 1395, Opus 4.8 1615. The v2 suffix appears only on this page (Opus 4.8 page shows GDPval-AA without version label, Sonnet 5 at 1615 vs Opus 4.8 page 1890 for Opus) - version differences flagged, do not merge silently.
未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Safety evaluations section · row: Firefox 147 exploit development · quote_snippet: Neither of the Sonnet models could successfully develop a working exploit (both scored 0.0%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: firefox-147-exploit not yet in data/benchmarks/.