Claude 3.5 Sonnet
Anthropic · 2024-06-20 · 类别未确认
Claude 3.5 Sonnet
Anthropic 旗舰中杯模型,在代码生成、多步复杂推理与视觉理解上全面超越上一代 Opus 与同代旗舰,是智能体与工程编程的行业标杆。
- 输入模态
- 文本 / 图像
- 上下文
- 200K tokens
- 参数
- 未公开
- 价格(每百万 tokens)
- 输入 3 / 输出 15
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。
未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier intelligence at 2x the speed · table: Claude 3.5 Sonnet benchmarks · row: Graduate level reasoning - GPQA, Diamond · figure: 官方页图 cf2c754458e9102b7334731fb18a965bfeb7ad08-2200x1894.png(benchmark 总表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro | Llama-400b (early snapshot)) · quote_snippet: Claude 3.5 Sonnet sets new industry benchmarks for graduate-level reasoning (GPQA), undergraduate-level knowledge (MMLU), and coding proficiency (HumanEval).
{
"harness": null,
"tools": null,
"shots": "0-shot CoT",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表(2026-09-01,Claude 3.5 Sonnet 列)。表脚注:5-shot CoT GPQA with maj@32 得 67.2%。同行对比:Claude 3 Opus 50.4%(0-shot CoT)、GPT-4o 53.6%(0-shot CoT)、Gemini 1.5 Pro / Llama-400b 未报。原 pending 行补值升级。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier intelligence at 2x the speed · table: Claude 3.5 Sonnet benchmarks · row: Undergraduate level knowledge - MMLU · figure: 官方页图 cf2c754458e9102b7334731fb18a965bfeb7ad08-2200x1894.png(benchmark 总表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro | Llama-400b (early snapshot)) · quote_snippet: Claude 3.5 Sonnet sets new industry benchmarks for graduate-level reasoning (GPQA), undergraduate-level knowledge (MMLU), and coding proficiency (HumanEval).
{
"harness": null,
"tools": null,
"shots": "5-shot(主值);0-shot CoT(次值 88.3)",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表(2026-09-01,Claude 3.5 Sonnet 列)。表内两行协议:5-shot 88.7%(标 **)、0-shot CoT 88.3%;表脚注:5-shot CoT prompting 得 90.4%。同行对比(5-shot):Claude 3 Opus 86.8%、Gemini 1.5 Pro 85.9%、Llama-400b 86.1%、GPT-4o 未报;(0-shot CoT):Opus 85.7%、GPT-4o 88.7%(绿标最高)。原 pending 行补值升级。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier intelligence at 2x the speed · table: Claude 3.5 Sonnet benchmarks · row: Code - HumanEval · figure: 官方页图 cf2c754458e9102b7334731fb18a965bfeb7ad08-2200x1894.png(benchmark 总表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro | Llama-400b (early snapshot)) · quote_snippet: Claude 3.5 Sonnet sets new industry benchmarks for graduate-level reasoning (GPQA), undergraduate-level knowledge (MMLU), and coding proficiency (HumanEval).
{
"harness": null,
"tools": null,
"shots": "0-shot",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表(2026-09-01,Claude 3.5 Sonnet 列)。同行对比(均 0-shot):Claude 3 Opus 84.9%、GPT-4o 90.2%、Gemini 1.5 Pro 84.1%、Llama-400b 84.1%。原 pending 行补值升级。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier intelligence at 2x the speed · table: Claude 3.5 Sonnet benchmarks · row: Multilingual math - MGSM · figure: 官方页图 cf2c754458e9102b7334731fb18a965bfeb7ad08-2200x1894.png(benchmark 总表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro | Llama-400b (early snapshot)) · quote_snippet: Claude 3.5 Sonnet operates at twice the speed of Claude 3 Opus.
{
"harness": null,
"tools": null,
"shots": "0-shot CoT",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表(2026-09-01,Claude 3.5 Sonnet 列)。同行对比:Claude 3 Opus 90.7%(0-shot CoT)、GPT-4o 90.5%(0-shot CoT)、Gemini 1.5 Pro 87.5%(8-shot)、Llama-400b 未报。新行:官方总表补录。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier intelligence at 2x the speed · table: Claude 3.5 Sonnet benchmarks · row: Reasoning over text - DROP, F1 score · figure: 官方页图 cf2c754458e9102b7334731fb18a965bfeb7ad08-2200x1894.png(benchmark 总表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro | Llama-400b (early snapshot))
{
"harness": null,
"tools": null,
"shots": "3-shot",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表(2026-09-01,Claude 3.5 Sonnet 列)。同行对比:Claude 3 Opus 83.1(3-shot)、GPT-4o 83.4(3-shot)、Gemini 1.5 Pro 74.9(variable shots)、Llama-400b 83.5(3-shot,pre-trained model)。新行:官方总表补录。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier intelligence at 2x the speed · table: Claude 3.5 Sonnet benchmarks · row: Mixed evaluations - BIG-Bench-Hard · figure: 官方页图 cf2c754458e9102b7334731fb18a965bfeb7ad08-2200x1894.png(benchmark 总表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro | Llama-400b (early snapshot))
{
"harness": null,
"tools": null,
"shots": "3-shot CoT",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表(2026-09-01,Claude 3.5 Sonnet 列)。同行对比:Claude 3 Opus 86.8%(3-shot CoT)、Gemini 1.5 Pro 89.2%(3-shot CoT)、Llama-400b 85.3%(3-shot CoT,pre-trained model)、GPT-4o 未报。新行:官方总表补录。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier intelligence at 2x the speed · table: Claude 3.5 Sonnet benchmarks · row: Math problem-solving - MATH · figure: 官方页图 cf2c754458e9102b7334731fb18a965bfeb7ad08-2200x1894.png(benchmark 总表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro | Llama-400b (early snapshot))
{
"harness": null,
"tools": null,
"shots": "0-shot CoT",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表(2026-09-01,Claude 3.5 Sonnet 列)。此行非全场最高:GPT-4o 76.6%(0-shot CoT,绿标)。其余对比:Claude 3 Opus 60.1%(0-shot CoT)、Gemini 1.5 Pro 67.7%(4-shot)、Llama-400b 57.8%(4-shot CoT)。新行:官方总表补录。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier intelligence at 2x the speed · table: Claude 3.5 Sonnet benchmarks · row: Grade school math - GSM8K · figure: 官方页图 cf2c754458e9102b7334731fb18a965bfeb7ad08-2200x1894.png(benchmark 总表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro | Llama-400b (early snapshot))
{
"harness": null,
"tools": null,
"shots": "0-shot CoT",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表(2026-09-01,Claude 3.5 Sonnet 列),数值 96.4 与原迁移行一致。同行对比:Claude 3 Opus 95.0%(0-shot CoT)、Gemini 1.5 Pro 90.8%(11-shot)、Llama-400b 94.1%(8-shot CoT)、GPT-4o 未报。原 pending 行补定位升级。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art vision · table: Claude 3.5 Sonnet vision evals · row: Visual math reasoning - MathVista (testmini) · figure: 官方页图 caff3d60763b27b59fe33e4ae984530f0dba4ddb-2200x1110.png(vision evals 表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro) · quote_snippet: Claude 3.5 Sonnet is our strongest vision model yet, surpassing Claude 3 Opus on standard vision benchmarks.
{
"harness": null,
"tools": null,
"shots": "0-shot CoT",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页 vision 图表(2026-09-01,Claude 3.5 Sonnet 列)。同行对比(均 0-shot CoT):Claude 3 Opus 50.5%、GPT-4o 63.8%、Gemini 1.5 Pro 63.9%。新行:官方 vision 表补录。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art vision · table: Claude 3.5 Sonnet vision evals · row: Visual question answering - MMMU (val) · figure: 官方页图 caff3d60763b27b59fe33e4ae984530f0dba4ddb-2200x1110.png(vision evals 表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro) · quote_snippet: Claude 3.5 Sonnet is our strongest vision model yet, surpassing Claude 3 Opus on standard vision benchmarks.
{
"harness": null,
"tools": null,
"shots": "0-shot CoT",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页 vision 图表(2026-09-01,Claude 3.5 Sonnet 列)。此行非全场最高:GPT-4o 69.1%(0-shot CoT,绿标)。其余对比:Claude 3 Opus 59.4%、Gemini 1.5 Pro 62.2%(均 0-shot CoT)。新行:官方 vision 表补录。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art vision · table: Claude 3.5 Sonnet vision evals · row: Chart Q&A - Relaxed accuracy (test) · figure: 官方页图 caff3d60763b27b59fe33e4ae984530f0dba4ddb-2200x1110.png(vision evals 表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro) · quote_snippet: These step-change improvements are most noticeable for tasks that require visual reasoning, like interpreting charts and graphs.
{
"harness": null,
"tools": null,
"shots": "0-shot CoT",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页 vision 图表(2026-09-01,Claude 3.5 Sonnet 列)。同行对比(均 0-shot CoT):Claude 3 Opus 80.8%、GPT-4o 85.7%、Gemini 1.5 Pro 87.2%。新行:官方 vision 表补录。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art vision · table: Claude 3.5 Sonnet vision evals · row: Document visual Q&A - ANLS score, test · figure: 官方页图 caff3d60763b27b59fe33e4ae984530f0dba4ddb-2200x1110.png(vision evals 表;列: Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro) · quote_snippet: Claude 3.5 Sonnet can also accurately transcribe text from imperfect images.
{
"harness": null,
"tools": null,
"shots": "0-shot",
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页 vision 图表(2026-09-01,Claude 3.5 Sonnet 列)。同行对比(均 0-shot):Claude 3 Opus 89.3%、GPT-4o 92.8%、Gemini 1.5 Pro 93.1%。新行:官方 vision 表补录。