Claude Haiku 4.5
Anthropic · 2025-10-15 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Claude Haiku 4.5
发布文将其定位为小而快的轻量档模型,编码性能接近 Sonnet 4 而成本仅约三分之一、速度快一倍以上,计算机使用能力超过 Sonnet 4。11 项评测覆盖智能体编码、工具调用、计算机使用与数学推理,亮点为 AIME 2025(带 Python 工具)96.3% 与 τ2-bench Telecom 83.0%。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 1 / 输出 5
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Opening chart + Benchmarks table · row: Agentic coding - SWE-bench Verified · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 gives you similar levels of coding performance [to Sonnet 4] but at one-third the cost and more than twice the speed
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。prose makes only the qualitative "similar to Sonnet 4 coding performance" claim; numbers are chart-only. Vision-assisted read (notes only): Haiku 4.5 73.3 vs Sonnet 4.5 77.2, Sonnet 4 72.7, GPT-5 72.8 (GPT-5 high) / 74.5 (GPT-5-Codex), Gemini 2.5 Pro 67.2. Needs human chart confirmation. 2026-09-01 补充:prose Methodology 段亦印出 73.3%(50 trials 平均、无测试时计算、128K thinking budget、全量 500 题),与图表值一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks (comparison table) · row: Agentic terminal coding - Terminal-Bench · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 is one of our most powerful models to date. See footnotes for methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。no numeric score appears in prose; values live in chart images. Vision-assisted read (notes only, per goal.md 12.5 OCR cannot flip status): Haiku 4.5 41.0 vs Sonnet 4.5 50.0, Sonnet 4 36.4, GPT-5 43.8, Gemini 2.5 Pro 25.3. Version not stated on page. Needs human chart confirmation. 2026-09-01 补充:prose Methodology 段印出 Terminal-Bench 平均过程 40.21%(6 次 no thinking)与 41.75%(5 次 32K thinking),表值 41.0 为报告均值。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks (comparison table) · row: Agentic tool use - t2-bench (Retail) · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 is one of our most powerful models to date. See footnotes for methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。no numeric score appears in prose; values live in chart images. Vision-assisted read (notes only, per goal.md 12.5 OCR cannot flip status): Haiku 4.5 83.2 vs Sonnet 4.5 86.2, Sonnet 4 83.8, GPT-5 81.1, Gemini 2.5 Pro not reported. new-benchmark id tau2-bench (distinct from existing tau-bench). Needs human chart confirmation.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks (comparison table) · row: Agentic tool use - t2-bench (Airline) · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 is one of our most powerful models to date. See footnotes for methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。no numeric score appears in prose; values live in chart images. Vision-assisted read (notes only, per goal.md 12.5 OCR cannot flip status): Haiku 4.5 63.6 vs Sonnet 4.5 70.0, Sonnet 4 63.0, GPT-5 62.6. new-benchmark id tau2-bench. Needs human chart confirmation.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks (comparison table) · row: Agentic tool use - t2-bench (Telecom) · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 is one of our most powerful models to date. See footnotes for methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。no numeric score appears in prose; values live in chart images. Vision-assisted read (notes only, per goal.md 12.5 OCR cannot flip status): Haiku 4.5 83.0 vs Sonnet 4.5 98.0, Sonnet 4 49.6, GPT-5 96.7. new-benchmark id tau2-bench. Needs human chart confirmation.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks (comparison table) · row: Computer use - OSWorld · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 is one of our most powerful models to date. See footnotes for methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。no numeric score appears in prose; values live in chart images. Vision-assisted read (notes only, per goal.md 12.5 OCR cannot flip status): Haiku 4.5 50.7 vs Sonnet 4.5 61.4 (cross-checks exactly with the Sonnet 4.5 release prose), Sonnet 4 42.2, GPT-5 and Gemini 2.5 Pro not reported. Prose separately claims Haiku 4.5 surpasses Sonnet 4 at certain tasks like using computers. Needs human chart confirmation.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks (comparison table) · row: High school math competition - AIME 2025 (python) · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 is one of our most powerful models to date. See footnotes for methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。no numeric score appears in prose; values live in chart images. Vision-assisted read (notes only, per goal.md 12.5 OCR cannot flip status): Haiku 4.5 96.3 vs Sonnet 4.5 100, Sonnet 4 70.5, GPT-5 99.6 (cross-checks exactly with the OpenAI GPT-5 page), Gemini 2.5 Pro 88.0. Needs human chart confirmation.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks (comparison table) · row: High school math competition - AIME 2025 (no tools) · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 is one of our most powerful models to date. See footnotes for methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。no numeric score appears in prose; values live in chart images. Vision-assisted read (notes only, per goal.md 12.5 OCR cannot flip status): Haiku 4.5 80.7 vs Sonnet 4.5 87.0, GPT-5 94.6 (cross-checks exactly with the OpenAI GPT-5 page); Sonnet 4 and Gemini 2.5 Pro blank. Needs human chart confirmation.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks (comparison table) · row: Graduate-level reasoning - GPQA Diamond · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 is one of our most powerful models to date. See footnotes for methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。no numeric score appears in prose; values live in chart images. Vision-assisted read (notes only, per goal.md 12.5 OCR cannot flip status): Haiku 4.5 73.0 vs Sonnet 4.5 83.4, Sonnet 4 76.1, GPT-5 85.7 (cross-checks exactly with the OpenAI GPT-5 page), Gemini 2.5 Pro 86.4. Needs human chart confirmation.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks (comparison table) · row: Multilingual Q&A - MMMLU · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 is one of our most powerful models to date. See footnotes for methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。no numeric score appears in prose; values live in chart images. Vision-assisted read (notes only, per goal.md 12.5 OCR cannot flip status): Haiku 4.5 83.0 vs Sonnet 4.5 89.1, Sonnet 4 86.5, GPT-5 89.4, Gemini 2.5 Pro not reported. new-benchmark: mmmlu (Multilingual MMLU) not yet in data/benchmarks/, distinct from mmlu. Needs human chart confirmation.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks (comparison table) · row: Visual reasoning - MMMU (validation) · figure: images/02.webp(归档自官方页的 benchmark 对比总表 1920x1625;列: Claude Sonnet 4.5 | Claude Haiku 4.5 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro) · quote_snippet: Claude Haiku 4.5 is one of our most powerful models to date. See footnotes for methodology.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自官方页图表 029af67124b67bdf0b50691a8921b46252c023d2-1920x1625.png(2026-09-01 live 页独立复核,Claude Haiku 4.5 列)。no numeric score appears in prose; values live in chart images. Vision-assisted read (notes only, per goal.md 12.5 OCR cannot flip status): Haiku 4.5 73.2 vs Sonnet 4.5 77.8, Sonnet 4 74.4, GPT-5 84.2 (cross-checks exactly with the OpenAI GPT-5 page), Gemini 2.5 Pro 82.0. Needs human chart confirmation.