← 模型目录

Claude Opus 5.5

Anthropic · 2026-09-22 · 通用模型

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Claude Opus 5.5

面向代码、计算机操作和知识工作的模型。官方表格提供不同任务的分数,但协议与回退模型影响比较范围。

输入模态
官方资料未说明
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 4 / 输出 20

本变体的评测证据

terminalbench 66.4% 模型 claude-opus-5-5 · 版本 Terminal-Bench 4.0 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance and cost-effectiveness · table: Performance comparison · row: Terminal-Bench 4.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "adaptive thinking, xhigh",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。 Opus 5.5 标准误 ±2.6 个百分点;不是置信区间。未把公开榜的 5 trials/Claude Code 设置推定为本页设置。

打开官方来源

frontier-code 54.4% 模型 claude-opus-5-5 · 版本 FrontierCode v1.1 (Main) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance and cost-effectiveness · table: Performance comparison · row: FrontierCode v1.1 (Main)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "adaptive thinking, max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。

打开官方来源

cursor-bench 57.8% 模型 claude-opus-5-5 · 版本 CursorBench 4.0 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance and cost-effectiveness · table: Performance comparison · row: CursorBench 4.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "adaptive thinking, max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。

打开官方来源

gdpval-aa 1846 模型 claude-opus-5-5 · 版本 GDPval-AA v2.1 · 指标 elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance and cost-effectiveness · table: Performance comparison · row: GDPval-AA v2.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "adaptive thinking, max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。

打开官方来源

automationbench 40.0% 模型 claude-opus-5-5 · 版本 AutomationBench · 指标 pass_rate · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance and cost-effectiveness · table: Performance comparison · row: AutomationBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "adaptive thinking, max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Zapier 运行并报告,未启用回退模型,安全策略介入算失败;不同于其他行。

打开官方来源

hlehle 67.7% 模型 claude-opus-5-5 · 版本 Humanity's Last Exam with tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance and cost-effectiveness · table: Performance comparison · row: Humanity's Last Exam with tools

{
  "harness": null,
  "tools": [
    "enabled; tool list not disclosed"
  ],
  "shots": null,
  "reasoning_effort": "adaptive thinking, max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。

打开官方来源

terminal-bench-science 58.7% 模型 claude-opus-5-5 · 版本 Terminal-Bench-Science 0.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance and cost-effectiveness · table: Performance comparison · row: Terminal-Bench-Science 0.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "adaptive thinking, max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。 官方给各模型标准误范围 ±3.5–5 个百分点,未单列本模型。

打开官方来源

osworld 81.8% 模型 claude-opus-5-5 · 版本 OSWorld 2.0 partial · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance and cost-effectiveness · table: Performance comparison · row: OSWorld 2.0 partial

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "adaptive thinking, max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。

打开官方来源

chartography 89.0% 模型 claude-opus-5-5 · 版本 Chartography with tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance and cost-effectiveness · table: Performance comparison · row: Chartography with tools

{
  "harness": null,
  "tools": [
    "enabled; tool list not disclosed"
  ],
  "shots": null,
  "reasoning_effort": "adaptive thinking, max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。

打开官方来源

frontier-code 54.6% 模型 claude-opus-5-5 · 版本 FrontierCode v1.1 (Main), medium · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding · row: FrontierCode v1.1 (Main), medium

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "adaptive thinking, medium",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。

打开官方来源

cursor-bench 52.5% 模型 claude-opus-5-5 · 版本 CursorBench 4.0, medium · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Coding · row: CursorBench 4.0, medium

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "adaptive thinking, medium",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。

打开官方来源