← 模型目录

Claude Opus 4.6

Anthropic · 2026-02-05 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Claude Opus 4.6

Claude Opus 4.6 是 Anthropic 的旗舰模型,引入多档推理努力(low/medium/high/max)、自适应思考、上下文压缩与 1M 上下文 beta。评测覆盖代理编码/工具/搜索、法律与多语言问答:SWE-bench Verified 80.8、GPQA Diamond 91.3、ARC-AGI-2 68.8、BrowseComp 84.0。

输入模态
文本 / 图像
上下文
1M(beta)
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 10 / 输出 37.5 · 发布文给出的 premium 1M 上下文(beta)档定价;标准档定价发布文未给出

本变体的评测证据

terminalbench 该来源尚无可读数值 模型 claude-opus-4-6 · 版本 2.0 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening section · row: Terminal-Bench 2.0 · figure: bar-chart image b8cfd7ebd6c82febce5f428f519d68a5dcf5d16f-3840x2160.png; benchmark table image f9564dd2f758237bd9dbe775674c4a375aff1e8a-2600x2968.png (columns: Opus 4.6 / Opus 4.5 / Sonnet 4.5 / Gemini 3 Pro / GPT-5.2 all models); not machine-read · quote_snippet: it achieves the highest score on the agentic coding evaluation Terminal-Bench 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose claims highest score; absolute value renders in charts. Vision-assisted read (NOT machine-readable): table shows Opus 4.6 65.4% vs Opus 4.5 59.8%, Sonnet 4.5 51.0%, Gemini 3 Pro 56.2% (54.2% self-reported), GPT-5.2 64.7% (64.0% self-reported, Codex CLI). Status verified for the claim with score_status not_extracted (Opus 5 precedent).

打开官方来源

hlehle 40.0 (without tools) / 53.0 (with tools) 模型 claude-opus-4-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening section · row: Humanity's Last Exam · figure: images/07.webp(归档 benchmark 总表 2600x2968;列: Opus 4.6 | Opus 4.5 | Sonnet 4.5 | Gemini 3 Pro | GPT-5.2 (all models)) · quote_snippet: leads all other frontier models on Humanity's Last Exam, a complex multidisciplinary reasoning test

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp(2026-09-01,Opus 4.6 列,与先前 vision-read 逐格一致)。Vision-assisted read (notes only): Opus 4.6 40.0% without tools / 53.0% with tools vs Opus 4.5 30.8/43.4, Sonnet 4.5 17.7/33.6, Gemini 3 Pro 37.5/45.8, GPT-5.2 (Pro) 36.6/50.0. Claim verified in prose; numbers require human chart confirmation. 页内不一致留痕:总表 with tools 为 53.0,柱状图 images/05.webp 同格为 53.1,两官方图相差 0.1。原 text/verified 行补值。

打开官方来源

gdpval 该来源尚无可读数值 模型 claude-opus-4-6 · 版本 AA (Elo) · 指标 elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening section · row: GDPval-AA · figure: bar-chart image 6e29759b50e8b3a8363b38b1f573d854df968671-3840x2160.png; benchmark table image f9564dd2f758237bd9dbe775674c4a375aff1e8a-2600x2968.png (columns: Opus 4.6 / Opus 4.5 / Sonnet 4.5 / Gemini 3 Pro / GPT-5.2 all models); not machine-read · quote_snippet: On GDPval-AA [...] Opus 4.6 outperforms the industry's next-best model (OpenAI's GPT-5.2) by around 144 Elo points, and its own predecessor (Claude Opus 4.5) by 190 points

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

gdpval id already introduced by a prior batch; reused with variant AA (Elo scoring). Prose gives only Elo deltas (+144 vs GPT-5.2, +190 vs Opus 4.5). Vision-assisted table read (notes only): Opus 4.6 1606 / GPT-5.2 1462 / Opus 4.5 1416 / Sonnet 4.5 1277 / Gemini 3 Pro 1195 - the OCR deltas match the prose deltas exactly (1606-1462=144, 1606-1416=190), strong corroboration, but absolute values still need human chart confirmation.

打开官方来源

browsecomp 84.0 模型 claude-opus-4-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening section · row: BrowseComp · figure: images/07.webp 总表与 images/03.webp 柱状图(Agentic search / BrowseComp)双源一致。原 text/verified 行补值。 · quote_snippet: Opus 4.6 also performs better than any other model on BrowseComp, which measures a model's ability to locate hard-to-find information online

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp 总表与 images/03.webp 柱状图(2026-09-01,Opus 4.6 列,与先前 vision-read 逐格一致)。browsecomp id already introduced by a prior batch; reused. Vision-assisted read (notes only): Opus 4.6 84.0% vs Opus 4.5 67.8%, Sonnet 4.5 43.9%, Gemini 3 Pro 59.2% (Deep Research), GPT-5.2 77.9% (Pro). Claim verified in prose; numbers require human chart confirmation.

打开官方来源

mrcr 76% 模型 claude-opus-4-6 · 版本 v2 8-needle 1M · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 · row: MRCR v2, 8-needle 1M variant · quote_snippet: on the 8-needle 1M variant of MRCR v2 [...] Opus 4.6 scores 76%, whereas Sonnet 4.5 scores just 18.5%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mrcr not yet in data/benchmarks/ (OpenAI MRCR v2 needle-in-haystack family; cross-vendor note: the OpenAI GPT-5.5 page reports the same v2 8-needle family by token band, 512K-1M GPT-5.5 74.0%).

打开官方来源

biglaw-bench 90.2% 模型 claude-opus-4-6 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: First impressions (Harvey customer quote) · row: BigLaw Bench · quote_snippet: Claude Opus 4.6 achieved the highest BigLaw Bench score of any Claude model at 90.2%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: biglaw-bench not yet in data/benchmarks/. Score reported by Harvey (customer) in a quote on the official release page, not by Anthropic evals - attribution third_party_reported. Quote adds: 40% perfect scores and 84% above 0.8.

打开官方来源

swebench 80.8 模型 claude-opus-4-6 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 (benchmark table) · row: Agentic coding - SWE-bench Verified · figure: images/07.webp(归档 benchmark 总表 2600x2968;列: Opus 4.6 | Opus 4.5 | Sonnet 4.5 | Gemini 3 Pro | GPT-5.2 (all models)) · quote_snippet: The table below shows how Claude Opus 4.6 compares to our previous models and to other industry models on a variety of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp(2026-09-01,Opus 4.6 列,与先前 vision-read 逐格一致)。Opus 4.6 80.8 vs Opus 4.5 80.9 (predecessor marginally higher), Sonnet 4.5 77.2, Gemini 3 Pro 76.2, GPT-5.2 80.0.

打开官方来源

osworld 72.7 模型 claude-opus-4-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 (benchmark table) · row: Agentic computer use - OSWorld · figure: images/07.webp(归档 benchmark 总表 2600x2968;列: Opus 4.6 | Opus 4.5 | Sonnet 4.5 | Gemini 3 Pro | GPT-5.2 (all models)) · quote_snippet: The table below shows how Claude Opus 4.6 compares

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp(2026-09-01,Opus 4.6 列,与先前 vision-read 逐格一致)。Opus 4.6 72.7, Opus 4.5 66.3, Sonnet 4.5 61.4 (cross-checks exactly with the Sonnet 4.5 release prose), Gemini 3 Pro and GPT-5.2 not reported.

打开官方来源

tau2-bench Retail 91.9 / Telecom 99.3 模型 claude-opus-4-6 · 版本 Retail / Telecom split · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 (benchmark table) · row: Agentic tool use - t2-bench · figure: images/07.webp(归档 benchmark 总表 2600x2968;列: Opus 4.6 | Opus 4.5 | Sonnet 4.5 | Gemini 3 Pro | GPT-5.2 (all models)) · quote_snippet: The table below shows how Claude Opus 4.6 compares

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp(2026-09-01,Opus 4.6 列,与先前 vision-read 逐格一致)。PENDING: new-benchmark id tau2-bench. Vision-assisted read (notes only): Opus 4.6 Retail 91.9 / Telecom 99.3; Opus 4.5 88.9/98.2; Sonnet 4.5 86.2/98.0; Gemini 3 Pro 85.3/98.0; GPT-5.2 82.0/98.7. 单边行容纳 Retail/Telecom 两值:value 取 Retail,完整原文见 display。

打开官方来源

mcp-atlas 59.5 模型 claude-opus-4-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 (benchmark table) · row: Scaled tool use - MCP Atlas · figure: images/07.webp(归档 benchmark 总表 2600x2968;列: Opus 4.6 | Opus 4.5 | Sonnet 4.5 | Gemini 3 Pro | GPT-5.2 (all models)) · quote_snippet: The table below shows how Claude Opus 4.6 compares

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp(2026-09-01,Opus 4.6 列,与先前 vision-read 逐格一致)。PENDING: mcp-atlas id already introduced by a prior batch; reused. Vision-assisted read (notes only): Opus 4.6 59.5, Opus 4.5 62.3 (predecessor higher), Sonnet 4.5 43.8, Gemini 3 Pro 54.1, GPT-5.2 60.6.

打开官方来源

finance-agent 60.7 模型 claude-opus-4-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 (benchmark table) · row: Agentic financial analysis - Finance Agent · figure: images/07.webp(归档 benchmark 总表 2600x2968;列: Opus 4.6 | Opus 4.5 | Sonnet 4.5 | Gemini 3 Pro | GPT-5.2 (all models)) · quote_snippet: The table below shows how Claude Opus 4.6 compares

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp(2026-09-01,Opus 4.6 列,与先前 vision-read 逐格一致)。PENDING: finance-agent id already introduced by a prior batch; reused. Vision-assisted read (notes only): Opus 4.6 60.7, Opus 4.5 55.9, Sonnet 4.5 54.2, Gemini 3 Pro 44.1, GPT-5.2 56.6 (labeled 5.1).

打开官方来源

arc-agi 68.8 模型 claude-opus-4-6 · 版本 2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 (benchmark table) · row: Novel problem-solving - ARC AGI 2 · figure: images/07.webp(归档 benchmark 总表 2600x2968;列: Opus 4.6 | Opus 4.5 | Sonnet 4.5 | Gemini 3 Pro | GPT-5.2 (all models)) · quote_snippet: The table below shows how Claude Opus 4.6 compares

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp(2026-09-01,Opus 4.6 列,与先前 vision-read 一致)。Opus 4.6 68.8, Opus 4.5 37.6, Sonnet 4.5 13.6, Gemini 3 Pro 45.1 (Deep Thinking), GPT-5.2 54.2 (Pro).

打开官方来源

gpqa 91.3 模型 claude-opus-4-6 · 版本 Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 (benchmark table) · row: Graduate-level reasoning - GPQA Diamond · figure: images/07.webp(归档 benchmark 总表 2600x2968;列: Opus 4.6 | Opus 4.5 | Sonnet 4.5 | Gemini 3 Pro | GPT-5.2 (all models)) · quote_snippet: The table below shows how Claude Opus 4.6 compares

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp(2026-09-01,Opus 4.6 列,与先前 vision-read 逐格一致)。Opus 4.6 91.3, Opus 4.5 87.0, Sonnet 4.5 83.4, Gemini 3 Pro 91.9, GPT-5.2 93.2 (Pro).

打开官方来源

mmmu-pro 73.9 (without tools) / 77.3 (with tools) 模型 claude-opus-4-6 · 版本 without/with tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 (benchmark table) · row: Visual reasoning - MMMU Pro · figure: images/07.webp(归档 benchmark 总表 2600x2968;列: Opus 4.6 | Opus 4.5 | Sonnet 4.5 | Gemini 3 Pro | GPT-5.2 (all models)) · quote_snippet: The table below shows how Claude Opus 4.6 compares

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp(2026-09-01,Opus 4.6 列,与先前 vision-read 逐格一致)。Opus 4.6 73.9 no tools / 77.3 with tools; Opus 4.5 70.6/73.9; Sonnet 4.5 63.4/68.9; Gemini 3 Pro 81.0/not reported; GPT-5.2 79.5/80.4. new-benchmark: mmmu-pro not yet in data/benchmarks/ (prior batches have not migrated it). 单边行容纳 no-tools/with-tools 两值:value 取 without tools,完整原文见 display。

打开官方来源

mmmlu 91.1 模型 claude-opus-4-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 (benchmark table) · row: Multilingual Q&A - MMMLU · figure: images/07.webp(归档 benchmark 总表 2600x2968;列: Opus 4.6 | Opus 4.5 | Sonnet 4.5 | Gemini 3 Pro | GPT-5.2 (all models)) · quote_snippet: The table below shows how Claude Opus 4.6 compares

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/07.webp(2026-09-01,Opus 4.6 列,与先前 vision-read 逐格一致)。PENDING: new-benchmark: mmmlu (Multilingual MMLU) not yet in data/benchmarks/. Vision-assisted read (notes only): Opus 4.6 91.1, Opus 4.5 90.8, Sonnet 4.5 89.5, Gemini 3 Pro 91.8, GPT-5.2 89.6.

打开官方来源

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

mrcr 18.5% 模型 claude-sonnet-4-5 · 版本 v2 8-needle 1M · 指标 accuracy · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluating Claude Opus 4.6 · row: MRCR v2, 8-needle 1M variant · quote_snippet: Opus 4.6 scores 76%, whereas Sonnet 4.5 scores just 18.5%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mrcr not yet in data/benchmarks/. Prior Anthropic model score cited on this page.

打开官方来源