← 模型目录

Claude Sonnet 4.5

Anthropic · 2025-09-29 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Claude Sonnet 4.5

Anthropic 发布 Claude Sonnet 4.5,作为面向编码、智能体与计算机使用场景的 Sonnet 新代,并同步推出 Claude Agent SDK。评测集中在编码与智能体领域,亮点如 SWE-bench Verified 77.2%(并行测试期计算 82.0%)、τ2-bench Telecom 98.0%。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 3 / 输出 15 · 与 Sonnet 4 同价

本变体的评测证据

swebench 77.2 (standard) / 82.0 (with parallel test-time compute) 模型 claude-sonnet-4-5 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier intelligence · row: SWE-bench Verified · figure: images/02.webp(归档 benchmark 总表 2600x2288;列: Claude Sonnet 4.5 | Claude Opus 4.1 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Sonnet 4.5 is state-of-the-art on the SWE-bench Verified evaluation, which measures real-world software coding abilities

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01):Sonnet 4.5 列 77.2 标准 / 82.0 parallel test-time compute。原 text/verified 行补值。Absolute score renders only in chart images. Vision-assisted chart read (NOT machine-readable; per goal.md 12.5 OCR cannot flip status): 77.2% standard, 82.0% with parallel test-time compute (table image), vs Opus 4.1 74.5/79.4, Sonnet 4 72.7/80.2, GPT-5 72.8 / GPT-5-Codex 74.5, Gemini 2.5 Pro 67.2. Status kept verified for the SOTA claim with score_status not_extracted (Opus 5 precedent).

打开官方来源

osworld 61.4% 模型 claude-sonnet-4-5 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier intelligence · row: OSWorld · quote_snippet: On OSWorld, a benchmark that tests AI models on real-world computer tasks, Sonnet 4.5 now leads at 61.4%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

terminalbench 50.0 模型 claude-sonnet-4-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier intelligence (benchmark table) · row: Agentic terminal coding - Terminal-Bench · figure: images/02.webp(归档 benchmark 总表 2600x2288;列: Claude Sonnet 4.5 | Claude Opus 4.1 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: The model also shows improved capabilities on a broad range of evaluations including reasoning and math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 4.5 列,与先前 vision-read 一致)。Sonnet 4.5 50.0%, Opus 4.1 46.5%, Sonnet 4 36.4%, GPT-5 43.8%, Gemini 2.5 Pro 25.3%. Version (1.x/2.x) not stated on page. Needs human chart confirmation before flipping to verified.

打开官方来源

tau2-bench Retail 86.2 / Airline 70.0 / Telecom 98.0 模型 claude-sonnet-4-5 · 版本 Retail / Airline / Telecom split · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier intelligence (benchmark table) · row: Agentic tool use - t2-bench · figure: images/02.webp(归档 benchmark 总表 2600x2288;列: Claude Sonnet 4.5 | Claude Opus 4.1 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: improved capabilities on a broad range of evaluations

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 4.5 列,与先前 vision-read 一致)。tau2-bench 已入 data/benchmarks/(distinct from tau-bench)。Vision-assisted read (notes only): Sonnet 4.5 Retail 86.2 / Airline 70.0 / Telecom 98.0; Opus 4.1 86.8/63.0/71.5; Sonnet 4 83.8/63.0/49.6; GPT-5 81.1/62.6/96.7; Gemini 2.5 Pro not reported.

打开官方来源

aime-25 100 (with python) / 87.0 (no tools) 模型 claude-sonnet-4-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier intelligence (benchmark table) · row: High school math competition - AIME 2025 · figure: images/02.webp(归档 benchmark 总表 2600x2288;列: Claude Sonnet 4.5 | Claude Opus 4.1 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: shows substantial gains in reasoning and math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 4.5 列,与先前 vision-read 一致)。Sonnet 4.5 100% with python / 87.0% no tools; Opus 4.1 78.0; Sonnet 4 70.5; GPT-5 99.6 with python / 94.6 no tools; Gemini 2.5 Pro 88.0. GPT-5 values cross-check exactly with the OpenAI GPT-5 release page.

打开官方来源

gpqa 83.4 模型 claude-sonnet-4-5 · 版本 Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier intelligence (benchmark table) · row: Graduate-level reasoning - GPQA Diamond · figure: images/02.webp(归档 benchmark 总表 2600x2288;列: Claude Sonnet 4.5 | Claude Opus 4.1 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: improved capabilities on a broad range of evaluations

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 4.5 列,与先前 vision-read 一致)。Sonnet 4.5 83.4; Opus 4.1 81.0; Sonnet 4 76.1; GPT-5 85.7; Gemini 2.5 Pro 86.4. GPT-5 85.7 cross-checks exactly with the OpenAI GPT-5 release page (no-tools with-thinking).

打开官方来源

mmmlu 89.1 模型 claude-sonnet-4-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier intelligence (benchmark table) · row: Multilingual Q&A - MMMLU · figure: images/02.webp(归档 benchmark 总表 2600x2288;列: Claude Sonnet 4.5 | Claude Opus 4.1 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: improved capabilities on a broad range of evaluations

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 4.5 列,与先前 vision-read 一致)。mmmlu 已入 data/benchmarks/(distinct from mmlu)。Vision-assisted read (notes only): Sonnet 4.5 89.1; Opus 4.1 89.5; Sonnet 4 86.5; GPT-5 89.4; Gemini 2.5 Pro not reported.

打开官方来源

mmmu 77.8 模型 claude-sonnet-4-5 · 版本 validation · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier intelligence (benchmark table) · row: Visual reasoning - MMMU (validation) · figure: images/02.webp(归档 benchmark 总表 2600x2288;列: Claude Sonnet 4.5 | Claude Opus 4.1 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: improved capabilities on a broad range of evaluations

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 4.5 列,与先前 vision-read 一致)。Sonnet 4.5 77.8; Opus 4.1 77.1; Sonnet 4 74.4; GPT-5 84.2; Gemini 2.5 Pro 82.0. GPT-5 84.2 cross-checks exactly with the OpenAI GPT-5 release page.

打开官方来源

finance-agent 55.3 模型 claude-sonnet-4-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier intelligence (benchmark table) · row: Financial analysis - Finance Agent · figure: images/02.webp(归档 benchmark 总表 2600x2288;列: Claude Sonnet 4.5 | Claude Opus 4.1 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: improved capabilities on a broad range of evaluations

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.webp(2026-09-01,Sonnet 4.5 列,与先前 vision-read 一致)。finance-agent 复用既有 registry id。Vision-assisted read (notes only): Sonnet 4.5 55.3; Opus 4.1 50.9; Sonnet 4 44.5; GPT-5 46.9; Gemini 2.5 Pro 29.4.

打开官方来源

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

osworld 42.2% 模型 claude-sonnet-4 · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier intelligence · row: OSWorld · quote_snippet: Just four months ago, Sonnet 4 held the lead at 42.2%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior-generation Anthropic model score cited on this page.

打开官方来源