Claude Opus 5.5
Anthropic · 2026-09-22 · 通用模型
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Claude Opus 5.5
面向代码、计算机操作和知识工作的模型。官方表格提供不同任务的分数,但协议与回退模型影响比较范围。
- 输入模态
- 官方资料未说明
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 4 / 输出 20
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · table: Performance comparison · row: Terminal-Bench 4.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "adaptive thinking, xhigh",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。 Opus 5.5 标准误 ±2.6 个百分点;不是置信区间。未把公开榜的 5 trials/Claude Code 设置推定为本页设置。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · table: Performance comparison · row: FrontierCode v1.1 (Main)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "adaptive thinking, max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · table: Performance comparison · row: CursorBench 4.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "adaptive thinking, max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · table: Performance comparison · row: GDPval-AA v2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "adaptive thinking, max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · table: Performance comparison · row: AutomationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "adaptive thinking, max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Zapier 运行并报告,未启用回退模型,安全策略介入算失败;不同于其他行。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · table: Performance comparison · row: Humanity's Last Exam with tools
{
"harness": null,
"tools": [
"enabled; tool list not disclosed"
],
"shots": null,
"reasoning_effort": "adaptive thinking, max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · table: Performance comparison · row: Terminal-Bench-Science 0.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "adaptive thinking, max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。 官方给各模型标准误范围 ±3.5–5 个百分点,未单列本模型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · table: Performance comparison · row: OSWorld 2.0 partial
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "adaptive thinking, max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance and cost-effectiveness · table: Performance comparison · row: Chartography with tools
{
"harness": null,
"tools": [
"enabled; tool list not disclosed"
],
"shots": null,
"reasoning_effort": "adaptive thinking, max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding · row: FrontierCode v1.1 (Main), medium
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "adaptive thinking, medium",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Coding · row: CursorBench 4.0, medium
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "adaptive thinking, medium",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 生产安全策略开启;安全策略介入时网络安全任务回退 Opus 4.8,生物/前沿模型开发任务回退 Opus 5。