Claude Opus 5
Anthropic · 2026-07-24 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Claude Opus 5
Anthropic 发布 Claude Opus 5,effort 分档 high/xhigh/max 并提供 Fast 模式;页面数值多为图表与对比性表述(含大量内部评测)。已收录 18 项评测覆盖前沿编码、Agent 操作与法律/健康/生物专业域:亮点 DeepSWE v1.1 68.8、BrowseComp 90.8。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · figure: images/02.webp(归档 benchmark 总表 2600x2578;列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol) · quote_snippet: on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8's performance
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 5 列)。原 verified 行由总表补值。new-benchmark: frontier-bench not yet in data/benchmarks.json (related: frontier-code and frontierswe are distinct benchmarks). Claim verified in prose; absolute score renders only in the chart image. Status kept verified for the claim with score_status not_extracted because the numeric value exists on-page in image form and requires human chart reading.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · quote_snippet: at max effort, the model performs within 0.5% of Fable 5's peak score, but at half the cost per task
{
"harness": "Cursor",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cursor-bench not yet in data/benchmarks.json. No absolute score printed; relative claim only. Page also claims Opus 5 achieves greater performance at a given cost than all other models on high, xhigh and max effort.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · figure: images/02.webp(归档 benchmark 总表 2600x2578;列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol) · quote_snippet: On ARC-AGI 3 ... Opus 5's score is three times as high as the next-best model
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 5 列)。原 verified 行由总表补值。Maps to existing benchmark arc-agi, variant 3 (ARC-AGI 3). Relative claim only; no absolute score printed.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · figure: images/02.webp(归档 benchmark 总表 2600x2578;列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol) · quote_snippet: Opus 5's pass rate is around 1.5x the next-best model for the same cost per task
{
"harness": "Zapier AutomationBench",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 5 列)。原 verified 行由总表补值。new-benchmark: automationbench not yet in data/benchmarks/. Page names it 'Zapier AutomationBench' - same benchmark family as the AutomationBench rows in GLM/Kimi/DeepSeek releases; subset/version not stated on this page, so do not merge scores across releases. Relative claim only; an early-access quote also cites a 100% single-task pass.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · figure: images/02.webp(归档 benchmark 总表 2600x2578;列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol) · quote_snippet: surpassing Fable 5's best result at just over a third of the cost
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 5 列)。原 verified 行由总表补值。Maps to existing benchmark osworld, variant 2.0 (computer use). Claim: outperforms every other model at any given cost. Relative claim only; no absolute score printed (GPT-5.6 page reports 62.6% for Sol on the same variant).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Alignment and safety / Safety · figure: Chart image b22d18a4d2003401f96f866effd9a40b5518c4c5-3840x2160.png (left: vulnerability identification, right: exploit development); values not machine-read · quote_snippet: close to Mythos 5 at identifying ... far behind ... developing exploits
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: oss-fuzz (Anthropic-developed cybersecurity evaluation based on OSS-Fuzz) not yet in data/benchmarks.json. Qualitative claim verified in prose and figure caption; numeric values in chart image. Mythos 5 is an Anthropic model, so this is an intra-family comparison.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Alignment and safety / Alignment · figure: Chart image 76d4af96516ffca2aceb4c1d0b0a83e2720d874b-3840x2160.png · quote_snippet: Opus 5 scores 2.3 on overall misaligned behavior, the lowest of our recent models
{
"harness": "automated behavioral audit (Anthropic internal)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: anthropic-behavioral-audit not yet in data/benchmarks.json. Internal safety evaluation, lower is better; scale not defined on page. Included for completeness of the safety section, not comparable to capability benchmarks.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · quote_snippet: scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: anthropic-life-sciences-internal (unnamed internal suite covering structural biology, organic chemistry, bioinformatics). Only deltas versus Opus 4.8 are reported: +10.2pp on organic chemistry (spectroscopy-to-structure), +7.7pp on protein variant-effect tasks. Absolute scores not published.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness(benchmark 总表) · row: Knowledge work - GDPval-AA v2 · figure: 官方页总表 a8fb4f77…-2600x2578(列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 live 页总表视觉转写补行:Opus 5 1861 vs Fable 5 1747、Opus 4.8 1593、GPT-5.6 Sol 1736(同表引用记于 notes)。与 sonnet-5 页 GDPval-AA v2(Sonnet 5 1618)同版本族。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness(benchmark 总表) · row: Agentic search - BrowseComp · figure: 官方页总表 a8fb4f77…-2600x2578(列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 live 页总表视觉转写补行:Opus 5 90.8 vs Fable 5 87.4、Opus 4.8 84.3、GPT-5.6 Sol 90.4(同表引用记于 notes)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness(benchmark 总表) · row: Multidisciplinary reasoning - Humanity's Last Exam (no tools) · figure: 官方页总表 a8fb4f77…-2600x2578(列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 live 页总表视觉转写补行:no tools,Opus 5 56.3 vs Fable 5 56.5(此格 Fable 5 略高,页内高亮)、Opus 4.8 49.8、GPT-5.6 Sol 未报。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness(benchmark 总表) · row: Multidisciplinary reasoning - Humanity's Last Exam (with tools) · figure: 官方页总表 a8fb4f77…-2600x2578(列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 live 页总表视觉转写补行:with tools,Opus 5 64.7 vs Fable 5 63.9、Opus 4.8 57.9、GPT-5.6 Sol 未报。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness(benchmark 总表) · row: Agentic coding - DeepSWE v1.1 · figure: 官方页总表 a8fb4f77…-2600x2578(列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 live 页总表视觉转写补行:Opus 5 68.8 vs Fable 5 69.7、Opus 4.8 59.0、GPT-5.6 Sol 72.7(GPT-5.6 Sol 最高,页内高亮;此行 Opus 5 非 SOTA)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness(benchmark 总表) · row: Agentic coding - FrontierCode v1.1, Main · figure: 官方页总表 a8fb4f77…-2600x2578(列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 live 页总表视觉转写补行:Opus 5 53.4 vs Fable 5 53.5(页内高亮)、Opus 4.8 46.5、GPT-5.6 Sol 47.5。frontier-code 与 frontier-bench 为不同 benchmark。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness(benchmark 总表) · row: Legal - Legal Agent Benchmark, Held-out · figure: 官方页总表 a8fb4f77…-2600x2578(列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 live 页总表视觉转写补行:Opus 5 11.7 vs Fable 5 列 13.3(页内高亮)、Opus 4.8 10.4、GPT-5.6 Sol 2.5。all-pass 口径与 opus-4-8 页 Harvey 引述的 ">10% overall" 同族。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness(benchmark 总表) · row: Health - HealthBench Professional · figure: 官方页总表 a8fb4f77…-2600x2578(列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 live 页总表视觉转写补行:Opus 5 59.8 vs Fable 5 列标注 Mythos 5 66.0(页内高亮)、Opus 4.8 57.4、GPT-5.6 Sol 60.5。注意 Fable 5 列在此行标注为 Mythos 5 数值。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness(benchmark 总表) · row: Biology - BioMysteryBench (hard) · figure: 官方页总表 a8fb4f77…-2600x2578(列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 live 页总表视觉转写补行:hard,Opus 5 49.4 vs Fable 5 列标注 Mythos 5 46.5、Opus 4.8 42.4、GPT-5.6 Sol 未报。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness(benchmark 总表) · row: Biology - BioMysteryBench (human solved) · figure: 官方页总表 a8fb4f77…-2600x2578(列: Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}2026-09-01 live 页总表视觉转写补行:human solved,Opus 5 90.1 vs Fable 5 列标注 Mythos 5 89.0、Opus 4.8 88.5、GPT-5.6 Sol 未报。