Claude Sonnet 5.5
Anthropic · 2026-09-28 · 通用模型
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Claude Sonnet 5.5
面向日常任务、缺陷修复与文档、幻灯片、表格制作的更快更省的 Sonnet 模型,官方称比 Sonnet 5 快 30% 以上、多数任务成本最多低 30%。收录的 9 项评测以智能体编码与知识工作为主,如 Terminal-Bench 4.0 70.6%、GDPval-AA v2.1 1844。
- 输入模态
- 官方资料未说明
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 2 / 输出 10 · 缓存读取 $0.20/百万 token
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark comparison grid · row: Terminal-Bench 4.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Sonnet 5 为 10.3%;Opus 5.5 为 66.4%(脚注 1:Opus 5.5 取 Xhigh 最高分);GPT-6 Sol 列为空。协议仅记录页面明示字段,未写即留空;不作跨协议排名。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark comparison grid · row: FrontierCode 1.1 (Main), Max
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Sonnet 5 为 42.4%;Opus 5.5 为 54.4%;GPT-6 Sol 为 49.3%。脚注 2:Sonnet 5.5 在 Max 档分数低于 Xhigh 档(Max 档更常触发代码评审技能并拆给多个子智能体,导致超时或超范围改动)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark comparison grid · row: FrontierCode 1.1 (Main), Xhigh
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Xhigh",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同一表格行内 Sonnet 5.5 的第二个推理强度;其余列与 Max 行相同(Sonnet 5 42.4%、Opus 5.5 54.4%、GPT-6 Sol 49.3%)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark comparison grid · row: CursorBench 4.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Sonnet 5 为 34.1%;Opus 5.5 为 57.8%;GPT-6 Sol 列为空。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark comparison grid · row: GDPval-AA v2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Sonnet 5 为 1449;Opus 5.5 为 1846;GPT-6 Sol 为 1487(脚注 4:OpenAI 近期修复了 GPT-6 Sol 的图像理解缺陷,Artificial Analysis 官方分数可能尚未更新)。脚注 3:Artificial Analysis 在预发布部署上测得,页面称可能低估 Sonnet 5.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark comparison grid · row: AA-Briefcase v1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}页面称这是一项新的长程知识工作基准。Sonnet 5 为 1359;Opus 5.5 为 1822;GPT-6 Sol 为 1483(脚注 4)。脚注 3 同 GDPval-AA。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark comparison grid · row: Humanity's Last Exam, with tools
{
"harness": null,
"tools": [
"with tools"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Sonnet 5 为 54.9%;Opus 5.5 为 67.7%;GPT-6 Sol 列为空。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark comparison grid · row: OSWorld 2.1, partial
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}页面标注 partial;Sonnet 5 为 57.0%;Opus 5.5 为 81.8%;GPT-6 Sol 列为空。版本为 2.1,与 Opus 5.5 发布页的 2.0 不同,不可直接合并比较。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark comparison grid · row: Chartography, no tools
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}页面标注 no tools;Sonnet 5 为 15.6%;Opus 5.5 为 64.4%;GPT-6 Sol 为 53.6%(脚注 4)。Chartography 由 Surge AI 测得。