← 模型目录

Claude Opus 4.7 / Claude Opus 4.7 (fast)

Anthropic · 2026-04-16 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Claude Opus 4.7

Anthropic 将 Claude Opus 4.7 定位为在高级软件工程上较 Opus 4.6 显著提升、最难任务上增益尤为明显的旗舰模型,同时大幅增强视觉分辨率;这也是首个按 Project Glasswing 分级防护方案发布的模型。已收录评测覆盖智能体编码与终端、长上下文检索、计算机操作、网络安全与多学科推理等领域,SWE-bench Verified 87.6、GPQA Diamond 94.2。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 5 / 输出 25 · 官方页明示定价与 Opus 4.6 持平:$5 每百万输入 tokens、$25 每百万输出 tokens

本变体的评测证据

finance-agent 64.4 模型 claude-opus-4-7 · 版本 v1.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Real-world work section · row: Finance Agent · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: its state-of-the-art score on the Finance Agent evaluation (see table above)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。finance-agent id already introduced by a prior batch; reused. Prose claims SOTA with the number in the table image. Vision-assisted read (notes only, per goal.md 12.5): Opus 4.7 64.4 vs Opus 4.6 60.1, GPT-5.4 (Pro) 61.5, Gemini 3.1 Pro 59.7 - the GPT columns cross-check exactly with the OpenAI GPT-5.5 page (61.5/59.7). 原 text/verified 行补值(Finance Agent v1.1)。

打开官方来源

gdpval 1753 模型 claude-opus-4-7 · 版本 AA · 指标 未说明 · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Real-world work section · row: GDPval-AA · figure: images/03.webp(Knowledge work / GDPval-AA 柱状图:Opus 4.7 1753 | Opus 4.6 1619 | GPT-5.4 1674 | Gemini 3.1 Pro 1314) · quote_snippet: Opus 4.7 is also state-of-the-art on GDPval-AA, a third-party evaluation of economically valuable knowledge work across finance, legal, and other domains

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/03.webp(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。gdpval id already introduced by a prior batch; reused with variant AA. Prose SOTA claim; the Opus 4.7 benchmark table image (vision-read) does not include a GDPval-AA row - the value may sit in one of the auxiliary pre-release chart images. The later Opus 4.8 page table lists Opus 4.7 GDPval-AA at 1753 Elo (cross-page reference, not this page). 原 text/verified 行补值;unit 由 percent 修正为 elo。

打开官方来源

biglaw-bench 90.9% 模型 claude-opus-4-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Early-access tester quotes (Harvey) · row: BigLaw Bench · quote_snippet: Claude Opus 4.7 demonstrates strong substantive accuracy on BigLaw Bench for Harvey, scoring 90.9% at high effort

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

biglaw-bench id already introduced in this batch via claude-opus-4-6.json. Customer-reported score in a quote on the official page, not an Anthropic eval.

打开官方来源

cursor-bench 70% 模型 claude-opus-4-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Early-access tester quotes (Cursor) · row: CursorBench · quote_snippet: On CursorBench, Opus 4.7 is a meaningful jump in capabilities, clearing 70% versus Opus 4.6 at 58%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

cursor-bench id already used by claude-opus-5.json (variant 3.2); variant unspecified here. Customer-reported score.

打开官方来源

swebench-pro 64.3 模型 claude-opus-4-7 · 版本 Public · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Agentic coding - SWE-bench Pro · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 64.3 vs Opus 4.6 53.4, GPT-5.4 57.7 (cross-checks exactly with the OpenAI GPT-5.4 page), Gemini 3.1 Pro 54.2, Mythos Preview 77.8.

打开官方来源

swebench 87.6 模型 claude-opus-4-7 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Agentic coding - SWE-bench Verified · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 87.6 vs Opus 4.6 80.8 (cross-checks exactly with the Opus 4.6 page table), Gemini 3.1 Pro 80.6, Mythos Preview 93.9, GPT-5.4 not reported.

打开官方来源

terminalbench 69.4 模型 claude-opus-4-7 · 版本 2.0 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Agentic terminal coding - Terminal-Bench 2.0 · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 69.4 vs Opus 4.6 65.4, GPT-5.4 75.1 (labeled self-reported harness; cross-checks exactly with the OpenAI GPT-5.4 page), Gemini 3.1 Pro 68.5, Mythos Preview 82.0.

打开官方来源

hlehle 46.9 模型 claude-opus-4-7 · 版本 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Multidisciplinary reasoning - Humanity's Last Exam (no tools) · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 46.9 vs Opus 4.6 40.0, GPT-5.4 (Pro) 42.7, Gemini 3.1 Pro 44.4, Mythos Preview 56.8 (GPT/Gemini values cross-check exactly with the OpenAI GPT-5.5 page).

打开官方来源

hlehle 54.7 模型 claude-opus-4-7 · 版本 with tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Multidisciplinary reasoning - Humanity's Last Exam (with tools) · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 54.7 vs Opus 4.6 53.3, GPT-5.4 (Pro) 58.7, Gemini 3.1 Pro 51.4, Mythos Preview 64.7.

打开官方来源

browsecomp 79.3 模型 claude-opus-4-7 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Agentic search - BrowseComp · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 79.3, Opus 4.6 83.7 (predecessor higher on this row), GPT-5.4 (Pro) 89.3 (cross-checks exactly with the OpenAI GPT-5.4 page), Gemini 3.1 Pro 85.9, Mythos Preview 86.9.

打开官方来源

mcp-atlas 77.3 模型 claude-opus-4-7 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Scaled tool use - MCP-Atlas · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 77.3 vs Opus 4.6 75.8, GPT-5.4 68.1, Gemini 3.1 Pro 73.9, Mythos Preview not reported. Cross-page discrepancy: the OpenAI GPT-5.5 page cites Opus 4.7 MCP Atlas at 79.1 (footnoted as Scale AI April 2026 update) - different Scale AI snapshot, do not reconcile silently.

打开官方来源

osworld 78.0 模型 claude-opus-4-7 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Agentic computer use - OSWorld-Verified · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 78.0 vs Opus 4.6 72.7, GPT-5.4 75.0 (cross-checks exactly with the OpenAI GPT-5.4 page), Mythos Preview 79.6, Gemini 3.1 Pro not reported.

打开官方来源

cybergym 73.1 模型 claude-opus-4-7 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Cybersecurity vulnerability reproduction - CyberGym · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 73.1 vs Opus 4.6 73.8 (predecessor marginally higher), GPT-5.4 66.3, Mythos Preview 83.1, Gemini not reported. cybergym id already introduced by a prior batch; reused.

打开官方来源

gpqa 94.2 模型 claude-opus-4-7 · 版本 Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Graduate-level reasoning - GPQA Diamond · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 94.2 vs Opus 4.6 91.3, GPT-5.4 (Pro) 94.4, Gemini 3.1 Pro 94.3, Mythos Preview 94.6 (GPT/Gemini values cross-check exactly with the OpenAI GPT-5.5 page).

打开官方来源

charxiv-reasoning 82.1 (no tools) / 91.0 (with tools) 模型 claude-opus-4-7 · 版本 no tools / with tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Visual reasoning - CharXiv Reasoning · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 82.1 no tools / 91.0 with tools vs Opus 4.6 69.1/84.7, Mythos Preview 86.1/93.2, GPT-5.4 and Gemini not reported. charxiv-reasoning id already introduced by a prior batch; reused. 单边行容纳 no-tools/with-tools 两值:value 取 no tools,完整原文见 display。

打开官方来源

mmmlu 91.5 模型 claude-opus-4-7 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Opening benchmark table · row: Multilingual Q&A - MMMLU · figure: images/02.png(归档 benchmark 总表 2600x2638;列: Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro | Mythos Preview) · quote_snippet: it shows better results than Opus 4.6 across a range of benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Opus 4.7 列,与先前 vision-read 逐格一致)。Opus 4.7 91.5 vs Opus 4.6 91.1, Gemini 3.1 Pro 92.6, GPT-5.4 and Mythos Preview not reported. mmmlu id already introduced in this batch via claude-sonnet-4-5.json pending rows.

打开官方来源

screenspot-pro 79.5 (no tools) / 87.6 (with tools) 模型 claude-opus-4-7 · 版本 High resolution · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual navigation(ScreenSpot-Pro 柱状图) · row: Opus 4.7 High resolution · figure: 官方页图 e97dffe5…-1920x1080(ScreenSpot-Pro,Without tools / With tools 分组柱)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 live 页图表视觉转写补行:Visual navigation / ScreenSpot-Pro,Opus 4.7 High resolution 79.5(无工具)/ 87.6(有工具)。同图 Low resolution:Opus 4.7 69.0/85.9;Opus 4.6 57.7/83.1。原 17 行无此评测项。

打开官方来源

screenspot-pro 69.0 (no tools) / 85.9 (with tools) 模型 claude-opus-4-7 · 版本 Low resolution · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual navigation(ScreenSpot-Pro 柱状图) · row: Opus 4.7 Low resolution · figure: 官方页图 e97dffe5…-1920x1080(Without tools / With tools 分组柱)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 live 页图表视觉转写补行。同图 Opus 4.6 Low resolution 57.7(无工具)/ 83.1(有工具)为同表引用,记于 notes。

打开官方来源

graphwalks Parents 1M 75.1 / BFS 1M 58.6 模型 claude-opus-4-7 · 版本 Parents 1M / BFS 1M · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Long-context reasoning(GraphWalks 柱状图) · row: Opus 4.7 · figure: 官方页图 186551e6…-1920x1080(Parents 1M / BFS 1M 分组柱)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 live 页图表视觉转写补行:Long-context reasoning / GraphWalks,Parents 1M Opus 4.7 75.1 vs Opus 4.6 71.1;BFS 1M Opus 4.7 58.6 vs Opus 4.6 41.2(同表引用记于 notes)。原 17 行无此评测项。

打开官方来源

vending-bench-2 $10,937 模型 claude-opus-4-7 · 版本 未说明 · 指标 money_balance · 单位 usd 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Long-term coherence(Vending-Bench 2 柱状图) · row: Opus 4.7 · figure: 官方页图 d6b08133…-1920x1080(Money balance 柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 live 页图表视觉转写补行:Long-term coherence / Vending-Bench 2 期末资金余额,Opus 4.7 $10,937 vs Opus 4.6 $8,018(同表引用记于 notes;余额越高越好)。原 17 行无此评测项。

打开官方来源

swebench-multilingual 80.5 模型 claude-opus-4-7 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding(SWE-bench Multilingual and Multimodal 柱状图) · row: Multilingual - Opus 4.7 · figure: 官方页图 34fc5568…-1920x1080(Multilingual / Multimodal 分组柱) · quote_snippet: SWE-bench Verified, Pro, and Multilingual: Our memorization screens flag a subset of problems in these SWE-bench evals. Excluding any problems that show signs of memorization, Opus 4.7's margin of improvement over Opus 4.6 holds.

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 live 页图表视觉转写补行:SWE-bench Multilingual,Opus 4.7 80.5 vs Opus 4.6 77.8(同表引用记于 notes)。Methodology 注明记忆筛查不改变 4.7 相对 4.6 的领先幅度。原 17 行无此评测项。

打开官方来源

swebench-mm 34.5 模型 claude-opus-4-7 · 版本 internal implementation · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding(SWE-bench Multilingual and Multimodal 柱状图) · row: Multimodal (internal implementation) - Opus 4.7 · figure: 官方页图 34fc5568…-1920x1080 · quote_snippet: SWE-bench Multimodal: We used an internal implementation for both Opus 4.7 and Opus 4.6. Scores are not directly comparable to public leaderboard scores.

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 live 页图表视觉转写补行:SWE-bench Multimodal(内部实现,官方注明与公开榜单不可直接比较),Opus 4.7 34.5 vs Opus 4.6 27.1(同表引用记于 notes)。原 17 行无此评测项。

打开官方来源

anthropic-structural-biology 74.0 模型 claude-opus-4-7 · 版本 Structural Biology · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Biomolecular reasoning(Structural Biology 柱状图) · row: Opus 4.7 · figure: 官方页图 925c2e9b…-1920x1080

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: anthropic-structural-biology(Anthropic Biomolecular reasoning / Structural Biology 内部评测)not yet in data/benchmarks/。2026-09-01 live 页图表视觉转写补行:Opus 4.7 74.0 vs Opus 4.6 30.9(同表引用记于 notes)。原 17 行无此评测项。

打开官方来源

Claude Opus 4.7 (fast)

本变体没有独立发布公告,OpenRouter 于 2026-05-12 上架 claude-opus-4.7-fast 是其公开痕迹;本次发布未报告该变体的评测数值。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

cursor-bench 58% 模型 claude-opus-4-6 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Early-access tester quotes (Cursor) · row: CursorBench · quote_snippet: clearing 70% versus Opus 4.6 at 58%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prior Anthropic model score quoted by Cursor on this page.

打开官方来源