Claude Opus 4 / Claude Sonnet 4
Anthropic · 2025-05-22 · 类别未确认
Claude Opus 4
Claude 4 发布中 Opus 4 为最强编码档,采用混合模式(近即时回答 + 扩展思考,扩展思考可与工具调用结合,beta)。已收录 22 项评测覆盖软件工程、终端、知识与 Agent:亮点 SWE-bench Verified 72.5%(并行测试时计算 79.4%)。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 15 / 输出 75 · 发布文明确的 per MTok 定价
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening section · row: SWE-bench · quote_snippet: Claude Opus 4 is our most powerful model yet and the best coding model in the world, leading on SWE-bench (72.5%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose says "SWE-bench"; the chart section caption reads "Claude 4 models lead on SWE-bench Verified", so the 72.5 figure is the Verified variant. Companion chart also shows 79.4% with parallel test-time compute (see next row notes).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model improvements · row: SWE-bench Verified · figure: images/02.webp(归档柱状图 Software engineering / SWE-bench verified:Opus 4 72.5(79.4 并行测试时计算)/ Sonnet 4 72.7(80.2)/ Sonnet 3.7 62.3(70.3)/ Codex-1 72.1 / o3 69.1 / GPT-4.1 54.6 / Gemini 2.5 Pro Preview 63.2) · quote_snippet: Claude 4 models lead on SWE-bench Verified, a benchmark for performance on real software engineering tasks. See appendix for more on methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01):Opus 4 并行测试时计算 79.4,与 prose 72.5 相互印证。Vision-assisted chart read (NOT machine-readable; per goal.md 12.5 OCR cannot flip status): Opus 4 72.5 base / 79.4 with parallel test-time compute; Sonnet 4 72.7 / 80.2; Sonnet 3.7 62.3 / 70.3; OpenAI Codex-1 72.1; OpenAI o3 69.1; OpenAI GPT-4.1 54.6; Gemini 2.5 Pro 63.2. Chart footnote "Preview (05-06)". The 72.5/72.7 bases match the prose values exactly.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening section · row: Terminal-bench · quote_snippet: the best coding model in the world, leading on SWE-bench (72.5%) and Terminal-bench (43.2%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Page does not state the Terminal-Bench version; left unspecified. 2026-09-01 CSV 逐格核对:Opus 4 43.2 / 50.0(并行测试时计算);脚注 2:同脚手架换非 Claude 智能体 39.2。同行:Sonnet 4 35.5/41.3、Sonnet 3.7 35.2、o3 30.2、GPT-4.1 30.3、Gemini 2.5 Pro 25.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Agentic terminal coding - Terminal-bench · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06) · quote_snippet: Claude Opus 4 ... leading on SWE-bench (72.5%) and Terminal-bench (43.2%).
{
"harness": "Claude Code(脚注 2:同脚手架换非 Claude 智能体得 39.2%/33.5%)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Opus 4 列):43.2 / 50.0(并行测试时计算,脚注 5)。同行:Sonnet 3.7 35.2、o3 30.2、GPT-4.1 30.3、Gemini 2.5 Pro 25.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Graduate-level reasoning - GPQA Diamond · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Opus 4 列):79.6 / 83.3(并行测试时计算)。Appendix 另载 w/o extended thinking 74.9。同行:Sonnet 3.7 78.2、o3 83.3、GPT-4.1 66.3、Gemini 2.5 Pro 83.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Graduate-level reasoning - GPQA Diamond · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Opus 4 列)。脚注 5:并行测试时计算 = 多序列采样后由内部评分模型选单一最优。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Agentic tool use - TAU-bench (Retail / Airline) · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": "prompt addendum to Airline/Retail Agent Policy; extended thinking with tool use; max steps 30 -> 100",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Opus 4 列):Retail 81.4、Airline 59.6(extended thinking with tool use;无 w/o extended thinking 结果,见 appendix)。同行 Retail:Sonnet 3.7 81.2、o3 70.4、GPT-4.1 68.0;Airline:Sonnet 3.7 58.4、o3 52.0、GPT-4.1 49.4;Gemini 未报。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Multilingual Q&A - MMMLU · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": "14 门非英语语言平均(脚注 3)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Opus 4 列):88.8(14 门非英语语言平均)。Appendix 另载 w/o extended thinking 87.4。同行:Sonnet 3.7 85.9、o3 88.8、GPT-4.1 83.7;Gemini 未报。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Visual reasoning - MMMU (validation) · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Opus 4 列):76.5。Appendix 另载 w/o extended thinking 73.7。同行:Sonnet 3.7 75.0、o3 82.9、GPT-4.1 74.8、Gemini 2.5 Pro 79.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: High school math competition - AIME 2025 · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": "0.95(脚注 4)",
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Opus 4 列):75.5 / 90.0(并行测试时计算,nucleus sampling top_p 0.95)。Appendix 另载 w/o extended thinking 33.9。同行:Sonnet 3.7 54.8、o3 88.9、Gemini 2.5 Pro 83.0;GPT-4.1 未报。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: High school math competition - AIME 2025 · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": "0.95(脚注 4)",
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Opus 4 列)。脚注 5:并行测试时计算 = 多序列采样后由内部评分模型选单一最优。
Claude Sonnet 4
Claude 4 发布中 Sonnet 4 为同代高效档,同样采用近即时 + 扩展思考混合模式(扩展思考与工具调用结合为 beta)。已收录评测覆盖软件工程与 Agent:亮点 SWE-bench Verified 72.7%(并行测试时计算 80.2%)。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 3 / 输出 15 · 发布文明确的 per MTok 定价
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening section · row: SWE-bench · quote_snippet: Claude Sonnet 4 significantly improves on Sonnet 3.7, excelling in coding with a state-of-the-art 72.7% on SWE-bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Same Verified-variant reasoning as the Opus row; chart-read parallel-compute value 80.2% recorded in the Opus parallel row notes. 2026-09-01 CSV 逐格核对:72.7 基线与 prose/chart 一致;80.2 并行测试时计算另立行。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Agentic coding - SWE-bench Verified · figure: images/02.webp(归档柱状图 Software engineering / SWE-bench verified,与 CSV 逐格一致) · quote_snippet: Claude 4 models lead on SWE-bench Verified, a benchmark for performance on real software engineering tasks.
{
"harness": "bash/editor tools scaffold",
"tools": "bash + file editing (string replacement)",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": "0.95",
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "10 trials averaged (footnote 1)",
"aggregation": null,
"judge": null
}视觉转写+CSV 逐格核对(2026-09-01,Sonnet 4 列):80.2 并行测试时计算;与 prose(appendix 高算力段 "79.4% and 80.2% for Opus 4 and Sonnet 4 respectively")一致。同行:Sonnet 3.7 62.3/70.3、o3 69.1、GPT-4.1 54.6、Gemini 2.5 Pro Preview 63.2。脚注 5:并行测试时计算 = 多序列采样后由内部评分模型选单一最优。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Agentic terminal coding - Terminal-bench · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06) · quote_snippet: Claude Opus 4 ... leading on SWE-bench (72.5%) and Terminal-bench (43.2%).
{
"harness": "Claude Code(脚注 2:同脚手架换非 Claude 智能体得 39.2%/33.5%)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Sonnet 4 列):35.5(Claude Code 智能体)/ 41.3(并行测试时计算)。脚注 2:换与非 Claude 模型相同智能体得 33.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Agentic terminal coding - Terminal-bench · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06) · quote_snippet: Claude Opus 4 ... leading on SWE-bench (72.5%) and Terminal-bench (43.2%).
{
"harness": "Claude Code(脚注 2:同脚手架换非 Claude 智能体得 39.2%/33.5%)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Sonnet 4 列)。脚注 5:并行测试时计算 = 多序列采样后由内部评分模型选单一最优。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Graduate-level reasoning - GPQA Diamond · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Sonnet 4 列):75.4 / 83.8(并行测试时计算)。Appendix 另载 w/o extended thinking 70.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Graduate-level reasoning - GPQA Diamond · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Sonnet 4 列)。脚注 5:并行测试时计算 = 多序列采样后由内部评分模型选单一最优。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Agentic tool use - TAU-bench (Retail / Airline) · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": "prompt addendum to Airline/Retail Agent Policy; extended thinking with tool use; max steps 30 -> 100",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Sonnet 4 列):Retail 80.5、Airline 60.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Multilingual Q&A - MMMLU · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": "14 门非英语语言平均(脚注 3)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Sonnet 4 列):86.5。Appendix 另载 w/o extended thinking 85.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: Visual reasoning - MMMU (validation) · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Sonnet 4 列):74.4。Appendix 另载 w/o extended thinking 72.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: High school math competition - AIME 2025 · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": "0.95(脚注 4)",
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Sonnet 4 列):70.5 / 85.0(并行测试时计算)。Appendix 另载 w/o extended thinking 33.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Claude 4(benchmarks table) · table: Benchmarks Table · row: High school math competition - AIME 2025 · figure: 官方页 benchmarksTable CSV(cdn.sanity.io/files/4zrzovbb/website/0827c45e02042cf4c6158ea6c04abefedf2b728a.csv;列: Opus 4 | Sonnet 4 | Sonnet 3.7 | o3 | GPT-4.1 | Gemini 2.5 Pro Preview 05-06)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": "0.95(脚注 4)",
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CSV 逐格核对(2026-09-01,Sonnet 4 列)。脚注 5:并行测试时计算 = 多序列采样后由内部评分模型选单一最优。