Claude Fable 5.1 / Claude Mythos 5.1
Anthropic · 2026-09-01 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Claude Fable 5.1
发布文定位:编码与知识工作上世界最先进的模型,其研究能力是 AI 将参与科学进步的早期一瞥。本档已收录 9 项基准,覆盖智能体科研、智能体编码、知识工作、计算机操作与商业工作流;亮点:Terminal-Bench 4.0 55.8%、Humanity's Last Exam(with tools)65.0%。
- 输入模态
- 文本
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 10 / 输出 50 · 与 Fable 5 持平;降价来自缓存读——缓存读 $0.25/MTok(较此前降 75%),典型负载较 Fable 5 降约 25%,重代理负载最高约 45%
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: A new performance frontier · table: 页面 DOM 对比表(BenchmarkGrid,整页唯一 <table>) · row: Agentic scientific research — Terminal-Bench-Science 0.1 [1]
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: Terminal-Bench-Science——实体已随本发布建档(data/benchmarks/terminal-bench-science.json)。竞品列(同表原文):Fable 5 24.7% / Opus 5 29.0% / GPT-5.6 Sol 22.4%。脚注[1]:标准误 ±3.5–4.5 pts/模型;公开榜(3 trials/task,Claude Code harness)印 Opus 5 30.0% / Fable 5 21.4%,Anthropic 自建 setup 复现为 29.0% / 24.7%(within noise)——表内数值为 Anthropic setup 口径。另有 Accuracy-vs-Cost SVG 图(effort low→max),逐点值未转录。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: A new performance frontier · table: 页面 DOM 对比表(BenchmarkGrid,整页唯一 <table>) · row: Agentic coding — Terminal-Bench 4.0(Fable 5.1 列)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同格印双值:55.8%(Fable 5.1)+ 60.9% (Mythos 5.1),Mythos 值另立一条边(--terminalbench-4-0-mythos)。竞品列:Fable 5 42.0% / Opus 5 52.3% / GPT-5.6 Sol 37.3%。Accuracy-vs-Cost 图脚注:Fable 5.1 与 Mythos 5.1 为同一底层模型,分差来自旧版 cyber 护栏干预的任务;护栏改进后官方预计两模型差距将显著缩小。跨页提示:Fable 5 在本页 Terminal-Bench 4.0 记 42.0%,其 2026-06-09 发布页记 Terminal-Bench 2.1 88.0%——版本不同(2.1 vs 4.0),不可直接比较。effort 曲线逐点值未转录。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: A new performance frontier · table: 页面 DOM 对比表(BenchmarkGrid,整页唯一 <table>) · row: Knowledge work — GDPval-AA v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}页面只印数值未印单位;unit=elo 取自 gdpval-aa 实体定义(ELO 制)。竞品列:Fable 5 1723 / Opus 5 1824 / GPT-5.6 Sol 1711。跨页冲突待裁定:同一模型 Fable 5 在其 2026-06-09 发布页记 GDPval-AA 1932(该页未标 v2),与本页 1723 差 209 分——疑为 GDPval-AA 版本快照不同,跨页比较前必须对齐 variant。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: A new performance frontier · table: 页面 DOM 对比表(BenchmarkGrid,整页唯一 <table>) · row: Computer use — OSWorld 2.0 [2](partial 行)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}竞品列:Fable 5 72.9% / Opus 5 75.4% / GPT-5.6 Sol 未报告(—)。脚注[2]:分数基于基准作者 2026-08 任务版;Fable 5 与 Opus 5 由 Anthropic 同条件重跑;任务文件与旧版不同,与已发表 OSWorld 2.0 数值不可直接比较(也是 GPT 列空缺的原因)——跨页比较前必须对齐任务版快照。评测条件:生产安全护栏开启,护栏干预任务上 Fable 5.1 与 Fable 5 均记 0 分,本值含该惩罚。partial / strict 为页面分行的两种判分口径。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: A new performance frontier · table: 页面 DOM 对比表(BenchmarkGrid,整页唯一 <table>) · row: Computer use — OSWorld 2.0(strict 行)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}竞品列:Fable 5 36.1% / Opus 5 39.6% / GPT-5.6 Sol 未报告(—)。条件与 partial 行相同(脚注[2] 2026-08 任务版不可比 + 生产护栏零分条款)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: A new performance frontier · table: 页面 DOM 对比表(BenchmarkGrid,整页唯一 <table>) · row: Multidisciplinary reasoning — Humanity's Last Exam(no tools 行)
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}无工具口径为页面行名明示(protocol.tools=[])。竞品列:Fable 5 57.8% / Opus 5 56.6% / GPT-5.6 Sol 未报告。另有 HLE Accuracy-vs-Cost SVG 图(no tools / with tools × effort low/max),逐点值未转录。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: A new performance frontier · table: 页面 DOM 对比表(BenchmarkGrid,整页唯一 <table>) · row: Multidisciplinary reasoning — Humanity's Last Exam(with tools 续行,benchmark 名留空)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}页面未列具体工具集,protocol.tools 保持 null(不臆测);无工具对照行见 --hlehle-no-tools。竞品列:Fable 5 63.8% / Opus 5 63.6% / GPT-5.6 Sol 未报告。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: A new performance frontier · table: 页面 DOM 对比表(BenchmarkGrid,整页唯一 <table>) · row: Business workflows — AutomationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}竞品列:Fable 5 17.1% / Opus 5 26.9% / GPT-5.6 Sol 19.6%。评测条件:生产护栏开启;护栏干预任务上页面只明示 Fable 5 记 0 分(未披露 Fable 5.1 在本行有护栏干预)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: A new performance frontier · table: 页面 DOM 对比表(BenchmarkGrid,整页唯一 <table>) · row: Agentic coding — CursorBench 3.2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}竞品列:Fable 5 70.5% / Opus 5 70.0% / GPT-5.6 Sol 67.2%。另有 CursorBench Accuracy-vs-Cost SVG 图(effort low→max),逐点值未转录。
Claude Mythos 5.1
发布文定位:与 Fable 5.1 同一模型、安全档不同,面向网络防御与生命科学专业工作的受信任访问档。本档已收录 1 项基准:Terminal-Bench 4.0 60.9%(官方脚注明示分差来自旧版 cyber 护栏干预的任务)。
- 输入模态
- 文本
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: A new performance frontier · table: 页面 DOM 对比表(BenchmarkGrid,整页唯一 <table>) · row: Agentic coding — Terminal-Bench 4.0(Fable 5.1 列内括注 60.9% (Mythos 5.1))
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Mythos 5.1 在本页唯一有数值的基准。值印于 Fable 5.1 列内括注(DOM 机读);Fable 5.1 同格 55.8% 另立一条边。差值来源(图脚注原文):两模型同一底层模型,分差反映旧版、精度较低的 cyber 护栏干预过的任务;护栏改进后预计差距显著缩小。竞品列 Fable 5 / Opus 5 / GPT-5.6 Sol 为各自模型数值,非 Mythos 同档对照。