Llama 4 Maverick / Llama 4 Scout / Llama 4 Behemoth
Meta / Llama · 2025-04-05 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Llama 4 Maverick
发布文称 Maverick 为 Llama 4 herd 中的主力多模态模型(17B 激活、128 专家、400B 总参,MoE 架构)。评测覆盖多模态理解、知识、代码与克丘亚语翻译,亮点为 DocVQA 94.4 与 ChartQA 90.0(LMArena Elo 1417 归属于实验性聊天版本)。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 400B-A17B MoE
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Takeaways · quote_snippet: an experimental chat version scoring ELO of 1417 on LMArena
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard",
"judge": null
}Maps to existing benchmark arena. The page itself qualifies the score as an EXPERIMENTAL CHAT VERSION — the headline Elo is not attributed to the base released weights; treat as a distinct arena submission snapshot when comparing against other vendors' flagship entries.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Maverick instruction-followed benchmarks) · row: Image Reasoning / MMMU · figure: images/03.png(归档 Llama 4 Maverick instruction-tuned benchmarks 表 1920x1638) · quote_snippet: Image Reasoning / MMMU
{
"harness": null,
"tools": null,
"shots": 1,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute; multiple generations averaged for high-variance benchmarks",
"judge": null
}视觉转写自归档图 images/03.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Table image visual reading Maverick: 73.4 (Gemini 2.0 Flash 71.7, GPT-4o 69.1; DeepSeek v3.1 no multimodal support). Footnote 1 (explicit on page): Llama results = 1-shot, temperature 0, no majority voting / parallel test-time compute. Footnote 2: non-Llama columns are highest available self-reported results, non-thinking models only.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Maverick instruction-followed benchmarks) · row: Image Reasoning / MathVista · figure: images/03.png(归档 Llama 4 Maverick instruction-tuned benchmarks 表 1920x1638) · quote_snippet: Image Reasoning / MathVista
{
"harness": null,
"tools": null,
"shots": 1,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/03.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Maverick: 73.7 (Gemini 2.0 Flash 73.1, GPT-4o 63.8).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Maverick instruction-followed benchmarks) · row: Image Understanding / ChartQA · figure: images/03.png(归档 Llama 4 Maverick instruction-tuned benchmarks 表 1920x1638) · quote_snippet: Image Understanding / ChartQA
{
"harness": null,
"tools": null,
"shots": 1,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/03.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Maverick: 90.0 (Gemini 2.0 Flash 88.3, GPT-4o 85.7).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Maverick instruction-followed benchmarks) · row: Image Understanding / DocVQA (test) · figure: images/03.png(归档 Llama 4 Maverick instruction-tuned benchmarks 表 1920x1638) · quote_snippet: DocVQA (test)
{
"harness": null,
"tools": null,
"shots": 1,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/03.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Maverick: 94.4 (GPT-4o 92.8; Gemini 2.0 Flash not reported).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Maverick instruction-followed benchmarks) · row: Coding / LiveCodeBench (10/01/2024-02/01/2025) · figure: images/03.png(归档 Llama 4 Maverick instruction-tuned benchmarks 表 1920x1638) · quote_snippet: LiveCodeBench (10/01/2024-02/01/2025)
{
"harness": null,
"tools": null,
"shots": 1,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute; multiple generations averaged (high-variance benchmark)",
"judge": null
}视觉转写自归档图 images/03.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Maverick: 43.4 (Gemini 2.0 Flash 34.5, GPT-4o 32.3 sourced from LCB leaderboard, DeepSeek v3.1 45.8 internal / 49.2 self-reported with unknown date range — footnote 3 flags the date-range mismatch). Same window label as Grok 3's LCB rows but Meta adds 1-shot/temp-0 conditions; still treat as vendor-specific protocol.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Maverick instruction-followed benchmarks) · row: Reasoning & Knowledge / MMLU Pro · figure: images/03.png(归档 Llama 4 Maverick instruction-tuned benchmarks 表 1920x1638) · quote_snippet: Reasoning & Knowledge / MMLU Pro
{
"harness": null,
"tools": null,
"shots": 1,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/03.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Maverick: 80.5 (Gemini 2.0 Flash 77.6, DeepSeek v3.1 81.2). Meta's 1-shot temp-0 condition differs from Llama 3.1's 5-shot CoT MMLU-Pro — protocol drift within the same vendor family.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Maverick instruction-followed benchmarks) · row: Reasoning & Knowledge / GPQA Diamond · figure: images/03.png(归档 Llama 4 Maverick instruction-tuned benchmarks 表 1920x1638) · quote_snippet: Reasoning & Knowledge / GPQA Diamond
{
"harness": null,
"tools": null,
"shots": 1,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute; multiple generations averaged (high-variance benchmark)",
"judge": null
}视觉转写自归档图 images/03.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Maverick: 69.8 (Gemini 2.0 Flash 60.1, DeepSeek v3.1 68.4, GPT-4o 53.6). Meta marks GPQA Diamond high-variance and averages multiple generations.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Maverick instruction-followed benchmarks) · row: Multilingual / Multilingual MMLU · figure: images/03.png(归档 Llama 4 Maverick instruction-tuned benchmarks 表 1920x1638) · quote_snippet: Multilingual / Multilingual MMLU
{
"harness": null,
"tools": null,
"shots": 1,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/03.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。new-benchmark: multilingual-mmlu not yet in data/benchmarks/ (distinct from global-mmlu-lite used by Google and from mmlu/mmlu-pro). Visual reading Maverick: 84.6 (GPT-4o 81.5).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Maverick instruction-followed benchmarks) · row: Long Context / MTOB (half book) eng → kgv/kgv → eng · figure: images/03.png(归档 Llama 4 Maverick instruction-tuned benchmarks 表 1920x1638) · quote_snippet: MTOB (half book) eng → kgv/kgv → eng
{
"harness": null,
"tools": null,
"shots": 1,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/03.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。new-benchmark: mtob not yet in data/benchmarks/ (Mission to Outer Babylon book-length translation benchmark). Visual reading Maverick: 54.0/46.4 (Gemini 2.0 Flash 48.4/39.8; DeepSeek v3.1 and GPT-4o limited to 128K context). Footnote 4: specialized long-context evals not traditionally reported for generalist models — Meta shares internal runs.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Maverick instruction-followed benchmarks) · row: Long Context / MTOB (full book) eng → kgv/kgv → eng · figure: images/03.png(归档 Llama 4 Maverick instruction-tuned benchmarks 表 1920x1638) · quote_snippet: MTOB (full book) eng → kgv/kgv → eng
{
"harness": null,
"tools": null,
"shots": 1,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/03.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Maverick: 50.8/46.7 (Gemini 2.0 Flash 45.5/39.6).
Llama 4 Scout
发布文称 Scout 以 17B 激活、16 专家、109B 总参实现单张 H100 可部署,并将上下文长度从 Llama 3 的 128K 提升至业界领先的 10M。评测以多模态与知识为主,亮点为 DocVQA 94.4 与 ChartQA 88.8。
- 输入模态
- 文本 / 图像
- 上下文
- 10M
- 参数
- 109B-A17B MoE
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Scout instruction-tuned benchmarks) · row: Image Reasoning / MMMU · figure: images/06.png(归档 Llama 4 Scout instruction-tuned benchmarks 表 1920x1359) · quote_snippet: Image Reasoning / MMMU
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute; multiple generations averaged for high-variance benchmarks",
"judge": null
}视觉转写自归档图 images/06.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Scout table footnote 1 differs from Maverick's: SCOUT rows are 0-SHOT (Maverick rows are 1-shot), both at temperature 0. Visual reading Scout: 69.4 (Llama 3.3 70B, Llama 3.1 405B, Gemma 3 27B, Mistral 3.1 24B, Gemini 2.0 Flash-Lite columns also in image).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Scout instruction-tuned benchmarks) · row: Image Reasoning / MathVista · figure: images/06.png(归档 Llama 4 Scout instruction-tuned benchmarks 表 1920x1359) · quote_snippet: Image Reasoning / MathVista
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/06.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Scout: 70.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Scout instruction-tuned benchmarks) · row: Image Understanding / ChartQA · figure: images/06.png(归档 Llama 4 Scout instruction-tuned benchmarks 表 1920x1359) · quote_snippet: Image Understanding / ChartQA
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/06.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Scout: 88.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Scout instruction-tuned benchmarks) · row: Image Understanding / DocVQA (test) · figure: images/06.png(归档 Llama 4 Scout instruction-tuned benchmarks 表 1920x1359) · quote_snippet: DocVQA (test)
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/06.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Scout: 94.4 (identical to Maverick's DocVQA figure in the visual read — plausible small-model tie worth re-checking on manual read).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Scout instruction-tuned benchmarks) · row: Coding / LiveCodeBench (10/01/2024-02/01/2025) · figure: images/06.png(归档 Llama 4 Scout instruction-tuned benchmarks 表 1920x1359) · quote_snippet: LiveCodeBench (10/01/2024-02/01/2025)
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute; multiple generations averaged (high-variance benchmark)",
"judge": null
}视觉转写自归档图 images/06.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Scout: 32.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Scout instruction-tuned benchmarks) · row: Reasoning & Knowledge / MMLU Pro · figure: images/06.png(归档 Llama 4 Scout instruction-tuned benchmarks 表 1920x1359) · quote_snippet: Reasoning & Knowledge / MMLU Pro
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/06.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Scout: 74.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Scout instruction-tuned benchmarks) · row: Reasoning & Knowledge / GPQA Diamond · figure: images/06.png(归档 Llama 4 Scout instruction-tuned benchmarks 表 1920x1359) · quote_snippet: Reasoning & Knowledge / GPQA Diamond
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute; multiple generations averaged (high-variance benchmark)",
"judge": null
}视觉转写自归档图 images/06.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Scout: 57.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Scout instruction-tuned benchmarks) · row: Long Context / MTOB (half book) eng → kgv/kgv → eng · figure: images/06.png(归档 Llama 4 Scout instruction-tuned benchmarks 表 1920x1359) · quote_snippet: MTOB (half book) eng → kgv/kgv → eng
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/06.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。new-benchmark: mtob (introduced in this batch by the Maverick rows). Visual reading Scout: 42.2/36.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Post-training our new models (Scout instruction-tuned benchmarks) · row: Long Context / MTOB (full book) eng → kgv/kgv → eng · figure: images/06.png(归档 Llama 4 Scout instruction-tuned benchmarks 表 1920x1359) · quote_snippet: MTOB (full book) eng → kgv/kgv → eng
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "no majority voting or parallel test-time compute",
"judge": null
}视觉转写自归档图 images/06.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Scout: 39.7/36.3.
Llama 4 Behemoth
发布文将其定位为仍在训练中的教师模型(约 2T 总参、288B 激活、16 专家),以当前最佳内部运行结果预览。评测显示知识与推理领先,亮点为 MATH-500 95.0 与 MMLU-Pro 82.2。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- ~2T-A288B MoE
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing the Llama 4 herd / We designed two efficient models · figure: images/07.png(归档 Llama 4 Behemoth instruction-tuned benchmarks 表 1920x1016) · quote_snippet: outperforms GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on STEM-focused benchmarks such as MATH-500
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/07.png(2026-09-01):Behemoth MATH 500 行 95.0(Claude Sonnet 3.7 82.2 / Gemini 2.0 Pro 91.8)。原 verified 行补值。Maps to existing benchmark math500. Prose names the benchmark with an outperforms claim but prints no number; Behemoth table image visually reads MATH-500 95.0 (Claude Sonnet 3.7 82.2, Gemini 2.0 Pro 91.8) — visual value kept in notes per §12.5. Behemoth is a still-training preview at announcement; its scores are internal-run snapshots by definition.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing the Llama 4 herd / We designed two efficient models · figure: images/07.png(归档 Llama 4 Behemoth instruction-tuned benchmarks 表 1920x1016) · quote_snippet: STEM-focused benchmarks such as MATH-500 and GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/07.png(2026-09-01):Behemoth GPQA Diamond 行 73.7(Claude Sonnet 3.7 68.0 / Gemini 2.0 Pro 64.7 / GPT-4.5 71.4)。原 verified 行补值。Prose names GPQA Diamond (same sentence as MATH-500) with an outperforms claim, no printed number. Behemoth table image visually reads GPQA Diamond 73.7 (Claude Sonnet 3.7 68.0, Gemini 2.0 Pro 64.7, GPT-4.5 71.4).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pushing Llama to new sizes: The 2T Behemoth (instruction-tuned benchmarks) · row: Coding / LiveCodeBench (10/01/2024-02/01/2025) · figure: images/07.png(归档 Llama 4 Behemoth instruction-tuned benchmarks 表 1920x1016) · quote_snippet: Coding / LiveCodeBench (10/01/2024-02/01/2025)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/07.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Behemoth table footnote 1: 'Llama model results represent our current best internal runs' — preview model, numbers are moving targets. Visual reading Behemoth: 49.4 (Gemini 2.0 Pro 36.0 sourced from LCB leaderboard; Claude Sonnet 3.7 / GPT-4.5 not reported).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pushing Llama to new sizes: The 2T Behemoth (instruction-tuned benchmarks) · row: Reasoning & Knowledge / MMLU Pro · figure: images/07.png(归档 Llama 4 Behemoth instruction-tuned benchmarks 表 1920x1016) · quote_snippet: Reasoning & Knowledge / MMLU Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/07.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Behemoth: 82.2 (Gemini 2.0 Pro 79.1).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pushing Llama to new sizes: The 2T Behemoth (instruction-tuned benchmarks) · row: Multilingual / Multilingual MMLU (OpenAI) · figure: images/07.png(归档 Llama 4 Behemoth instruction-tuned benchmarks 表 1920x1016) · quote_snippet: Multilingual MMLU (OpenAI)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/07.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。new-benchmark: multilingual-mmlu (introduced in this batch by the Maverick rows); Behemoth column labeled 'Multilingual MMLU (OpenAI)'. Visual reading Behemoth: 85.8 (Claude Sonnet 3.7 83.2, GPT-4.5 85.1).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pushing Llama to new sizes: The 2T Behemoth (instruction-tuned benchmarks) · row: Image Reasoning / MMMU · figure: images/07.png(归档 Llama 4 Behemoth instruction-tuned benchmarks 表 1920x1016) · quote_snippet: Image Reasoning / MMMU
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/07.png(2026-09-01,与先前 vision-read 一致;MTOB 为 chrF 双向对 eng→kgv/kgv→eng,value 取 eng→kgv 侧)。Visual reading Behemoth: 76.1 (Claude Sonnet 3.7 71.8, Gemini 2.0 Pro 72.7, GPT-4.5 74.4).