Claude Opus 4.8 / Claude Opus 4.8 (fast)
Anthropic · 2026-05-28 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Claude Opus 4.8
Anthropic 发布文将 Claude Opus 4.8 定位为 Opus 系列新旗舰,主打诚实性训练(放行自写代码缺陷的概率约为上代 1/4)与 effort 分档(默认 high,另有 extra/max),并随发布推出 Claude Code 动态工作流。已收录 9 项评测集中于编码、终端与计算机/知识工作 Agent:亮点 SWE-bench Pro 69.2、GDPval-AA Elo 1890。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 5 / 输出 25 · 发布文称定价与上一代持平(per MTok);另有 fast 模式 USD 10/50
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Early tester quotes (browser-agent partner) · row: Online-Mind2Web · quote_snippet: Claude Opus 4.8 is the strongest computer-use and browser-agent model we've tested, scoring 84% on Online-Mind2Web
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}online-mind2web id already introduced by a prior batch; reused. Customer-reported score in a quote on the official page; quote adds it is a meaningful jump over both Opus 4.7 and GPT-5.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Early tester quotes (Harvey) · row: Legal Agent Benchmark · quote_snippet: the highest score recorded on our Legal Agent Benchmark, and the first model to break 10% overall on the all-pass standard
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: legal-agent-benchmark not yet in data/benchmarks/. Customer (Harvey) benchmark; only the ">10% all-pass" threshold is printed, no absolute score, so value stays not_extracted. The Fable 5 page table later lists this benchmark with values (see cross-reference in claude-fable-5.json).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opus 4.8 capabilities (benchmark table) · row: Agentic coding - SWE-Bench Pro · figure: images/02.webp(归档 benchmark 总表 2600x1392;列: Opus 4.8 | Opus 4.7 | GPT-5.5 | Gemini 3.1 Pro) · quote_snippet: The table below shows how Opus 4.8 compares to its predecessor and to other models on tests of coding, agentic skills, reasoning, and practical knowledge work tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 4.8 列,与先前 vision-read 逐格一致)。Opus 4.8 69.2 vs Opus 4.7 64.3 (cross-checks exactly with the Opus 4.7 page table), GPT-5.5 58.6 (cross-checks exactly with the OpenAI GPT-5.5 page), Gemini 3.1 Pro 54.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opus 4.8 capabilities (benchmark table) · row: Agentic terminal coding - Terminal-Bench 2.1 · figure: images/02.webp(归档 benchmark 总表 2600x1392;列: Opus 4.8 | Opus 4.7 | GPT-5.5 | Gemini 3.1 Pro) · quote_snippet: The table below shows how Opus 4.8 compares to its predecessor and to other models on tests of coding, agentic skills, reasoning, and practical knowledge work tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 4.8 列,与先前 vision-read 逐格一致)。Opus 4.8 74.6 vs Opus 4.7 66.1, GPT-5.5 78.2, Gemini 3.1 Pro 70.3. Cross-page discrepancy: the Fable 5 and Sonnet 5 page tables both show Opus 4.8 at 82.7 on Terminal-Bench 2.1 and GPT-5.5 at 83.4 (Codex CLI) - either a later score revision or a harness difference; do not reconcile silently. 2026-09-01 live 页脚注补录:Terminal-Bench 2.1 全模型用 Terminus-2 公共 harness 报分;GPT-5.5 用 Codex CLI harness 的自报分为 83.4%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opus 4.8 capabilities (benchmark table) · row: Multidisciplinary reasoning - Humanity's Last Exam (no tools) · figure: images/02.webp(归档 benchmark 总表 2600x1392;列: Opus 4.8 | Opus 4.7 | GPT-5.5 | Gemini 3.1 Pro) · quote_snippet: The table below shows how Opus 4.8 compares to its predecessor and to other models on tests of coding, agentic skills, reasoning, and practical knowledge work tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 4.8 列,与先前 vision-read 逐格一致)。Opus 4.8 49.8 vs Opus 4.7 46.9, GPT-5.5 41.4, Gemini 3.1 Pro 44.4 (all cross-check exactly with the OpenAI GPT-5.5 and Opus 4.7 pages).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opus 4.8 capabilities (benchmark table) · row: Multidisciplinary reasoning - Humanity's Last Exam (with tools) · figure: images/02.webp(归档 benchmark 总表 2600x1392;列: Opus 4.8 | Opus 4.7 | GPT-5.5 | Gemini 3.1 Pro) · quote_snippet: The table below shows how Opus 4.8 compares to its predecessor and to other models on tests of coding, agentic skills, reasoning, and practical knowledge work tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 4.8 列,与先前 vision-read 逐格一致)。Opus 4.8 57.9 vs Opus 4.7 54.7, GPT-5.5 52.2, Gemini 3.1 Pro 51.4 (all cross-check exactly with prior pages).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opus 4.8 capabilities (benchmark table) · row: Agentic computer use - OSWorld-Verified · figure: images/02.webp(归档 benchmark 总表 2600x1392;列: Opus 4.8 | Opus 4.7 | GPT-5.5 | Gemini 3.1 Pro) · quote_snippet: The table below shows how Opus 4.8 compares to its predecessor and to other models on tests of coding, agentic skills, reasoning, and practical knowledge work tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 4.8 列,与先前 vision-read 逐格一致)。Opus 4.8 83.4 vs Opus 4.7 82.8, GPT-5.5 78.7, Gemini 3.1 Pro 76.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opus 4.8 capabilities (benchmark table) · row: Knowledge work - GDPval-AA · figure: images/02.webp(归档 benchmark 总表 2600x1392;列: Opus 4.8 | Opus 4.7 | GPT-5.5 | Gemini 3.1 Pro) · quote_snippet: The table below shows how Opus 4.8 compares to its predecessor and to other models on tests of coding, agentic skills, reasoning, and practical knowledge work tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 4.8 列,与先前 vision-read 逐格一致)。Opus 4.8 1890 vs Opus 4.7 1753, GPT-5.5 1769, Gemini 3.1 Pro 1314. Elo scoring; the Fable 5 page table lists Opus 4.8 at 1890 as well (consistent).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opus 4.8 capabilities (benchmark table) · row: Agentic financial analysis - Finance Agent v2 · figure: images/02.webp(归档 benchmark 总表 2600x1392;列: Opus 4.8 | Opus 4.7 | GPT-5.5 | Gemini 3.1 Pro) · quote_snippet: The table below shows how Opus 4.8 compares to its predecessor and to other models on tests of coding, agentic skills, reasoning, and practical knowledge work tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01,Opus 4.8 列,与先前 vision-read 逐格一致)。Opus 4.8 53.9 vs Opus 4.7 51.5, GPT-5.5 51.8, Gemini 3.1 Pro 43.0. Version label v2 per the table; distinct from the v1.1 rows on the Opus 4.7 page.
Claude Opus 4.8 (fast)
Claude Opus 4.8(fast)为同系 2.5 倍速度的快速档,定价 USD 10/50 每百万 tokens(发布文称约为历代 fast 模式的三分之一)。评测数值随主模型条目记录(如 Terminal-Bench 2.1 74.6),本条目无独立评测行。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 10 / 输出 50 · 2.5 倍速度的 fast 档,发布文称约为历代 fast 模式价格的三分之一
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。