Llama 3.1 405B / Llama 3.1 70B / Llama 3.1 8B
Meta / Llama · 2024-07-23 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Llama 3.1 405B
Meta 以「迄今最强模型」发布 Llama 3.1 系列,405B 为首个前沿级开源模型(dense decoder-only)。评测覆盖知识、数学、编码、多语言与长上下文,亮点如 MMLU 88.6%、GSM8K 96.8%。
- 输入模态
- 文本
- 上下文
- 128K
- 参数
- 405B dense
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: General / MMLU (0-shot, CoT) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: MMLU (5-shot, CoT)
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Benchmark table is a rendered image (fbcdn PNG, signed expiring URL — do not cite the URL as locator). Visual reading of Llama 3.1 405B column: 88.6 (Nemotron 4 78.7, GPT-4 85.4, GPT-4o 88.7, Claude 3.5 Sonnet 88.3). [2026-09-01 audit: image re-read live - label is "MMLU (0-shot, CoT)", shots corrected 5->0, stale pending-manual-read sentence removed.] Prose section 'Model evaluations' claims evaluation on 150+ benchmark datasets plus human evaluations, no printed numbers.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: General / MMLU PRO (5-shot, CoT) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: MMLU PRO (5-shot, CoT)
{
"harness": null,
"tools": null,
"shots": 5,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 73.3 (Nemotron 62.7, GPT-4 64.8, GPT-4o 74.0, Claude 3.5 Sonnet 77.0).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: General / IFEval · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: General / IFEval
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 88.6 (Nemotron 85.1, GPT-4 84.3, GPT-4o 85.6, Claude 3.5 Sonnet 88.0).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Code / HumanEval (0-shot) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: Code / HumanEval (0-shot)
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 89.0 (Nemotron 73.2, GPT-4 86.6, GPT-4o 90.2, Claude 3.5 Sonnet 92.0).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Code / MBPP EvalPlus (base) (0-shot) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: MBPP EvalPlus (base) (0-shot)
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: mbpp-evalplus not yet in data/benchmarks/ (EvalPlus-validated MBPP base; distinct from humaneval/humanevalplus). Visual reading 405B: 88.6 (Nemotron 72.8, GPT-4 83.6, GPT-4o 87.8, Claude 3.5 Sonnet 90.5). [2026-09-01 audit: live image label reads "MBPP EvalPlus (base, 0-shot)"; 2026-09-01 二次复核:归档原图 models/2024-07-23-llama-3-1/images/02.png 高清判定标签为 "MBPP EvalPlus (base) (0-shot)",此前 live 低清截图误读为 3-shot,已回滚为 0-shot.]
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Math / GSM8K (8-shot, CoT) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: GSM8K (8-shot, CoT)
{
"harness": null,
"tools": null,
"shots": 8,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 96.8 (Nemotron 92.3 0-shot, GPT-4 94.2, GPT-4o 96.1, Claude 3.5 Sonnet 96.4 0-shot — per-shot conditions differ across columns as marked in the image).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Math / MATH (0-shot, CoT) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: MATH (0-shot, CoT)
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: math (full Hendrycks MATH, 0-shot CoT) — also introduced by google/gemini-2-0 in this batch; distinct from existing math500. Visual reading 405B: 73.8 (Nemotron 41.1, GPT-4 64.5, GPT-4o 76.6, Claude 3.5 Sonnet 71.1).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Reasoning / ARC Challenge (0-shot) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: ARC Challenge (0-shot)
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: arc-challenge not yet in data/benchmarks/ — AllenAI ARC-Challenge, COMPLETELY UNRELATED to existing arc-agi (Cholak abstraction benchmark); ids must not be merged. Visual reading 405B: 96.9 (Nemotron 94.6, GPT-4 96.4, GPT-4o 96.7, Claude 3.5 Sonnet 96.7).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Reasoning / GPQA (8-shot, CoT) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: GPQA (8-shot, CoT)
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 51.1 (GPT-4 41.4, GPT-4o 53.6, Claude 3.5 Sonnet 59.4; Nemotron not evaluated). 8-shot CoT condition differs from most later vendors' 0-shot GPQA Diamond rows — comparability caution. [2026-09-01 audit: live image label reads "GPQA (0-shot, CoT)" on BOTH tables; shots corrected 8->0.]
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Tool use / BFCL · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: Tool use / BFCL
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 88.5 (Nemotron 86.5, GPT-4 88.3, GPT-4o 80.5, Claude 3.5 Sonnet 90.2). Maps to existing benchmark bfcl.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Tool use / Nexus · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: Tool use / Nexus
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: nexus not yet in data/benchmarks/ (Nexus function-calling tool-use benchmark). Visual reading 405B: 58.7 (GPT-4 50.3, GPT-4o 56.1, Claude 3.5 Sonnet 45.7).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Long context / ZeroSCROLLS/QuALITY · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: ZeroSCROLLS/QuALITY
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: zeroscrolls-quality not yet in data/benchmarks/. Visual reading 405B: 95.2 (GPT-4 95.2, GPT-4o 90.5, Claude 3.5 Sonnet 90.5 — tie with GPT-4 at top).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Long context / InfiniteBench/En.MC · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: InfiniteBench/En.MC
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: infinitebench-en-mc not yet in data/benchmarks/ (InfiniteBench English multiple-choice long-context task). Visual reading 405B: 83.4 (GPT-4 72.1, GPT-4o 82.5).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Long context / NIH/Multi-needle · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: NIH/Multi-needle
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Maps to existing benchmark niah (needle-in-a-haystack family), variant NIH/Multi-needle. Visual reading 405B: 98.1 (GPT-4 100.0, GPT-4o 100.0, Claude 3.5 Sonnet 90.8).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · row: Multilingual / Multilingual MGSM (0-shot) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: Multilingual MGSM (0-shot)
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Maps to existing benchmark mgsm. Visual reading 405B: 91.6 (GPT-4 85.9, GPT-4o 90.5, Claude 3.5 Sonnet 91.6 — tie at top).
Llama 3.1 70B
Llama 3.1 70B 为该系列中坚档:升级至 128K 上下文与多语言。亮点如 MMLU 86.0%、GSM8K 95.1%。
- 输入模态
- 文本
- 上下文
- 128K
- 参数
- 70B dense
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MMLU (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MMLU (0-shot, CoT) | 86
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 79.9 | GPT 3.5 Turbo 69.8。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MMLU PRO (5-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MMLU PRO (5-shot, CoT) | 66.4
{
"harness": null,
"tools": null,
"shots": 5,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 56.3 | GPT 3.5 Turbo 49.2。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: IFEval · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: IFEval | 87.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 72.7 | GPT 3.5 Turbo 69.9。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: HumanEval (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: HumanEval (0-shot) | 80.5
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 75.6 | GPT 3.5 Turbo 68.0。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MBPP EvalPlus (base, 3-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MBPP EvalPlus (base, 3-shot) | 86
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 78.6 | GPT 3.5 Turbo 82.0。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: GSM8K (8-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: GSM8K (8-shot, CoT) | 95.1
{
"harness": null,
"tools": null,
"shots": 8,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 88.2 | GPT 3.5 Turbo 81.6。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MATH (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MATH (0-shot, CoT) | 68
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 54.1 | GPT 3.5 Turbo 43.1。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: ARC Challenge (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: ARC Challenge (0-shot) | 94.8
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 88.7 | GPT 3.5 Turbo 83.7。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: GPQA (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: GPQA (0-shot, CoT) | 46.7
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 33.3 | GPT 3.5 Turbo 30.8。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: BFCL · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: BFCL | 84.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct – | GPT 3.5 Turbo 85.9。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: Nexus · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: Nexus | 56.7
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 48.5 | GPT 3.5 Turbo 37.2。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: ZeroSCROLLS/QuALITY · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: ZeroSCROLLS/QuALITY | 90.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct – | GPT 3.5 Turbo –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: InfiniteBench/En.MC · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: InfiniteBench/En.MC | 78.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct – | GPT 3.5 Turbo –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: NIH/Multi-needle · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: NIH/Multi-needle | 97.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct – | GPT 3.5 Turbo –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: Multilingual MGSM (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: Multilingual MGSM (0-shot) | 86.9
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 71.1 | GPT 3.5 Turbo 51.4。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
Llama 3.1 8B
Llama 3.1 8B 为该系列轻量档:升级至 128K 上下文与多语言。亮点如 GSM8K 84.5%、HumanEval 72.6%。
- 输入模态
- 文本
- 上下文
- 128K
- 参数
- 8B dense
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MMLU (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MMLU (0-shot, CoT) | 73
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 72.3(脚注: 5-shot, non-CoT) | Mistral 7B Instruct 60.5。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MMLU PRO (5-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MMLU PRO (5-shot, CoT) | 48.3
{
"harness": null,
"tools": null,
"shots": 5,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct 36.9。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: IFEval · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: IFEval | 80.4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 73.6 | Mistral 7B Instruct 57.6。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: HumanEval (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: HumanEval (0-shot) | 72.6
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 54.3 | Mistral 7B Instruct 40.2。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MBPP EvalPlus (base, 3-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MBPP EvalPlus (base, 3-shot) | 72.8
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 71.7 | Mistral 7B Instruct 49.5。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: GSM8K (8-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: GSM8K (8-shot, CoT) | 84.5
{
"harness": null,
"tools": null,
"shots": 8,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 76.7 | Mistral 7B Instruct 53.2。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MATH (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MATH (0-shot, CoT) | 51.9
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 44.3 | Mistral 7B Instruct 13.0。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: ARC Challenge (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: ARC Challenge (0-shot) | 83.4
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 87.6 | Mistral 7B Instruct 74.2。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: GPQA (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: GPQA (0-shot, CoT) | 32.8
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": "CoT",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct 28.8。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: BFCL · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: BFCL | 76.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct 60.4。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: Nexus · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: Nexus | 38.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 30.0 | Mistral 7B Instruct 24.7。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: ZeroSCROLLS/QuALITY · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: ZeroSCROLLS/QuALITY | 81
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: InfiniteBench/En.MC · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: InfiniteBench/En.MC | 65.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: NIH/Multi-needle · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: NIH/Multi-needle | 98.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: Multilingual MGSM (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: Multilingual MGSM (0-shot) | 68.9
{
"harness": null,
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 53.2 | Mistral 7B Instruct 29.9。与 405B 同名行共享 benchmark_id/variant,protocol 一致。