Claude Fable 5 / Claude Mythos 5
Anthropic · 2026-06-09 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Claude Fable 5
Anthropic 以「Claude Fable 5 与 Claude Mythos 5」联合公告发布 Fable 5:一款带安全护栏的 Mythos 级模型,在网络安全/生物化学/蒸馏话题上按分类器回退到 Opus 4.8(影响 <5% 会话)。评测覆盖智能体编码、知识工作、计算机使用与生物安全,亮点如 SWE-Bench Pro 80.3%、OSWorld-Verified 85.0%。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 10 / 输出 50 · 发布文记为低于 Mythos Preview 一半以上;订阅内含资格随时间调整
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Software engineering section · row: Cognition's FrontierCode evaluation · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: on Cognition's FrontierCode evaluation [...] Fable 5 scores highest among frontier models, even at medium effort
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01):FrontierCode (Diamond) xhigh 行 Mythos/Fable 列 29.3。new-benchmark: frontiercode not yet in data/benchmarks/ (Cognition eval for passing hard coding tasks at production-codebase standards). Naming inconsistency on the page: Anthropic prose calls it FrontierCode, Cognition's own quote calls it FrontierBench - same eval. Vision table read (notes only): FrontierCode (Diamond) xhigh Mythos/Fable 29.3 vs Opus 4.8 13.4, GPT 5.5 5.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Knowledge work section · row: Hebbia's Finance Benchmark · quote_snippet: On Hebbia's Finance Benchmark for senior-level reasoning, Fable 5 has the highest score of any model
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hebbia-finance-benchmark not yet in data/benchmarks/ (customer eval; score not printed in prose or read from a table row). Customer quote (Hebbia) also calls it the strongest finance-first model.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Agentic coding - SWE-Bench Pro · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos/Fable 80.3 vs Mythos Preview 77.8, Opus 4.8 69.2, GPT 5.5 58.6, Gemini 3.1 Pro 54.2 (Opus/GPT columns cross-check exactly with their own pages). Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Knowledge work - GDPval-AA · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos/Fable 1932 vs Opus 4.8 1890, GPT 5.5 1769, Gemini 3.1 Pro 1314. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Knowledge work vision - GDR.pdf · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos/Fable 29.8 vs Opus 4.8 22.5, GPT 5.5 24.9, Gemini 3.1 Pro 16.7. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries. new-benchmark: gdr-pdf not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Spatial reasoning - Blueprint-Bench 2 · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos/Fable 38.6 vs Opus 4.8 14.5, GPT 5.5 36.2, Gemini 3.1 Pro 26.5. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries. new-benchmark: blueprint-bench-2 not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Tool use - AutomationBench · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos/Fable 17.4 vs Opus 4.8 15.5, GPT 5.5 12.9, Gemini 3.1 Pro 9.6. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries. new-benchmark: automationbench not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Computer use - OSWorld-Verified · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos/Fable 85.0 vs Mythos Preview 85.4 (predecessor marginally higher), Opus 4.8 83.4, GPT 5.5 78.7, Gemini 3.1 Pro 76.2. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Legal - Legal Agent Benchmark · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos/Fable 13.3 vs Opus 4.8 10.4, GPT 5.5 2.1, Gemini 3.1 Pro 0.0. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries. legal-agent-benchmark id already introduced in this batch via claude-opus-4-8.json (Harvey).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Multidisciplinary reasoning - Humanity's Last Exam (no tools) · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos 59.0 / Fable (starred, safeguard-affected) vs Mythos Preview 56.8, Opus 4.8 49.8, GPT 5.5 41.4, Gemini 3.1 Pro 44.4. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Multidisciplinary reasoning - Humanity's Last Exam (with tools) · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos 64.5 / Fable (starred) vs Mythos Preview 64.7, Opus 4.8 57.9, GPT 5.5 52.2, Gemini 3.1 Pro 51.4. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Biology - BioMysteryBench (hard) · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos 46.1 / Fable (starred) vs Mythos Preview 29.6, Opus 4.8 40.0, GPT/Gemini not reported. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries. new-benchmark: biomysterybench not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Biology - BioMysteryBench (human solved) · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos 83.9 / Fable (starred) vs Mythos Preview 82.6, Opus 4.8 80.4. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries. new-benchmark: biomysterybench not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Agentic coding - Terminal-Bench 2.1 · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos 88.0 / Fable (starred) vs Opus 4.8 82.7, GPT 5.5 83.4 (Codex CLI), Gemini 3.1 Pro 70.7 (Gemini CLI). Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries. Cross-page discrepancy vs the Opus 4.8 own-page table (74.6/78.2) - see note there.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Cybersecurity - ExploitBench (Cap%) · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos 78.0 / Fable (starred) vs Mythos Preview 69.0, Opus 4.8 40.0, GPT 5.5 34.0, Gemini not reported. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries. exploitbench id already used by gpt-5-6.json; note the metric here is capability Cap% under Anthropic's methodology, not OpenAI's pass-rate framing - same benchmark family, different reporting.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening benchmark table · row: Health - HealthBench Professional · figure: images/02.webp(归档自官方页的 benchmark 总表,2600x2870;列: Claude Mythos 5 / Fable 5 | Claude Mythos Preview | Claude Opus 4.8 | GPT 5.5 | Gemini 3.1 Pro) · quote_snippet: The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.webp(2026-09-01 复核,数值与先前 vision-read 逐格一致)。Mythos 66.0 / Fable (starred) vs Mythos Preview 64.7, Opus 4.8 56.9, GPT 5.5 51.8. Table methodology footnote (vision-read): Mythos 5 and Fable 5 scores differ by 1-3pp; the table shows the higher of the two; starred rows differ more because Fable 5 falls back to Opus 4.8 on safeguard-triggered queries. healthbench id already introduced in this batch via gpt-5.json; Professional variant.
Claude Mythos 5
同公告中的 Claude Mythos 5 为与 Fable 5 同源的模型,在特定领域解除安全护栏,仅向 Project Glasswing 合作伙伴与后续可信访问计划开放,并报告药物设计环节约 10 倍加速等科研结果。评测与 Fable 5 共表(两模型分差 1–3 个百分点、取较高值),亮点如 Terminal-Bench 2.1 88.0%、GDPval-AA 1932 Elo。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。