← 模型目录

Llama 3.1 405B / Llama 3.1 70B / Llama 3.1 8B

Meta / Llama · 2024-07-23 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Llama 3.1 405B

Meta 以「迄今最强模型」发布 Llama 3.1 系列,405B 为首个前沿级开源模型(dense decoder-only)。评测覆盖知识、数学、编码、多语言与长上下文,亮点如 MMLU 88.6%、GSM8K 96.8%。

输入模态
文本
上下文
128K
参数
405B dense
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu 88.6 模型 llama-3-1-405b · 版本 0-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: General / MMLU (0-shot, CoT) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: MMLU (5-shot, CoT)

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Benchmark table is a rendered image (fbcdn PNG, signed expiring URL — do not cite the URL as locator). Visual reading of Llama 3.1 405B column: 88.6 (Nemotron 4 78.7, GPT-4 85.4, GPT-4o 88.7, Claude 3.5 Sonnet 88.3). [2026-09-01 audit: image re-read live - label is "MMLU (0-shot, CoT)", shots corrected 5->0, stale pending-manual-read sentence removed.] Prose section 'Model evaluations' claims evaluation on 150+ benchmark datasets plus human evaluations, no printed numbers.

打开官方来源

mmlu-pro 73.3 模型 llama-3-1-405b · 版本 5-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: General / MMLU PRO (5-shot, CoT) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: MMLU PRO (5-shot, CoT)

{
  "harness": null,
  "tools": null,
  "shots": 5,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 73.3 (Nemotron 62.7, GPT-4 64.8, GPT-4o 74.0, Claude 3.5 Sonnet 77.0).

打开官方来源

ifeval 88.6 模型 llama-3-1-405b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: General / IFEval · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: General / IFEval

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 88.6 (Nemotron 85.1, GPT-4 84.3, GPT-4o 85.6, Claude 3.5 Sonnet 88.0).

打开官方来源

humaneval 89.0 模型 llama-3-1-405b · 版本 0-shot · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Code / HumanEval (0-shot) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: Code / HumanEval (0-shot)

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 89.0 (Nemotron 73.2, GPT-4 86.6, GPT-4o 90.2, Claude 3.5 Sonnet 92.0).

打开官方来源

mbpp-evalplus 88.6 模型 llama-3-1-405b · 版本 base, 0-shot · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Code / MBPP EvalPlus (base) (0-shot) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: MBPP EvalPlus (base) (0-shot)

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: mbpp-evalplus not yet in data/benchmarks/ (EvalPlus-validated MBPP base; distinct from humaneval/humanevalplus). Visual reading 405B: 88.6 (Nemotron 72.8, GPT-4 83.6, GPT-4o 87.8, Claude 3.5 Sonnet 90.5). [2026-09-01 audit: live image label reads "MBPP EvalPlus (base, 0-shot)"; 2026-09-01 二次复核:归档原图 models/2024-07-23-llama-3-1/images/02.png 高清判定标签为 "MBPP EvalPlus (base) (0-shot)",此前 live 低清截图误读为 3-shot,已回滚为 0-shot.]

打开官方来源

gsm8k 96.8 模型 llama-3-1-405b · 版本 8-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Math / GSM8K (8-shot, CoT) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: GSM8K (8-shot, CoT)

{
  "harness": null,
  "tools": null,
  "shots": 8,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 96.8 (Nemotron 92.3 0-shot, GPT-4 94.2, GPT-4o 96.1, Claude 3.5 Sonnet 96.4 0-shot — per-shot conditions differ across columns as marked in the image).

打开官方来源

math 73.8 模型 llama-3-1-405b · 版本 0-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Math / MATH (0-shot, CoT) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: MATH (0-shot, CoT)

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: math (full Hendrycks MATH, 0-shot CoT) — also introduced by google/gemini-2-0 in this batch; distinct from existing math500. Visual reading 405B: 73.8 (Nemotron 41.1, GPT-4 64.5, GPT-4o 76.6, Claude 3.5 Sonnet 71.1).

打开官方来源

arc-challenge 96.9 模型 llama-3-1-405b · 版本 0-shot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Reasoning / ARC Challenge (0-shot) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: ARC Challenge (0-shot)

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: arc-challenge not yet in data/benchmarks/ — AllenAI ARC-Challenge, COMPLETELY UNRELATED to existing arc-agi (Cholak abstraction benchmark); ids must not be merged. Visual reading 405B: 96.9 (Nemotron 94.6, GPT-4 96.4, GPT-4o 96.7, Claude 3.5 Sonnet 96.7).

打开官方来源

gpqa 51.1 模型 llama-3-1-405b · 版本 0-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Reasoning / GPQA (8-shot, CoT) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: GPQA (8-shot, CoT)

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 51.1 (GPT-4 41.4, GPT-4o 53.6, Claude 3.5 Sonnet 59.4; Nemotron not evaluated). 8-shot CoT condition differs from most later vendors' 0-shot GPQA Diamond rows — comparability caution. [2026-09-01 audit: live image label reads "GPQA (0-shot, CoT)" on BOTH tables; shots corrected 8->0.]

打开官方来源

bfcl 88.5 模型 llama-3-1-405b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Tool use / BFCL · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: Tool use / BFCL

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Visual reading 405B: 88.5 (Nemotron 86.5, GPT-4 88.3, GPT-4o 80.5, Claude 3.5 Sonnet 90.2). Maps to existing benchmark bfcl.

打开官方来源

nexus 58.7 模型 llama-3-1-405b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Tool use / Nexus · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: Tool use / Nexus

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: nexus not yet in data/benchmarks/ (Nexus function-calling tool-use benchmark). Visual reading 405B: 58.7 (GPT-4 50.3, GPT-4o 56.1, Claude 3.5 Sonnet 45.7).

打开官方来源

zeroscrolls-quality 95.2 模型 llama-3-1-405b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Long context / ZeroSCROLLS/QuALITY · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: ZeroSCROLLS/QuALITY

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: zeroscrolls-quality not yet in data/benchmarks/. Visual reading 405B: 95.2 (GPT-4 95.2, GPT-4o 90.5, Claude 3.5 Sonnet 90.5 — tie with GPT-4 at top).

打开官方来源

infinitebench-en-mc 83.4 模型 llama-3-1-405b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Long context / InfiniteBench/En.MC · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: InfiniteBench/En.MC

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。new-benchmark: infinitebench-en-mc not yet in data/benchmarks/ (InfiniteBench English multiple-choice long-context task). Visual reading 405B: 83.4 (GPT-4 72.1, GPT-4o 82.5).

打开官方来源

niah 98.1 模型 llama-3-1-405b · 版本 NIH/Multi-needle · 指标 retrieval_accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Long context / NIH/Multi-needle · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: NIH/Multi-needle

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Maps to existing benchmark niah (needle-in-a-haystack family), variant NIH/Multi-needle. Visual reading 405B: 98.1 (GPT-4 100.0, GPT-4o 100.0, Claude 3.5 Sonnet 90.8).

打开官方来源

mgsm 91.6 模型 llama-3-1-405b · 版本 Multilingual, 0-shot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · row: Multilingual / Multilingual MGSM (0-shot) · figure: images/02.png(归档 405B benchmark 总表 3201x2217;列: Llama 3.1 405B | Nemotron 4 340B Instruct | GPT-4 (0125) | GPT-4 Omni | Claude 3.5 Sonnet) · quote_snippet: Multilingual MGSM (0-shot)

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自归档图 images/02.png(2026-09-01,Llama 3.1 405B 列,与先前 vision-read 一致;对照列 Nemotron 340B/GPT-4/GPT-4o/Claude 3.5 Sonnet 同表核对)。Maps to existing benchmark mgsm. Visual reading 405B: 91.6 (GPT-4 85.9, GPT-4o 90.5, Claude 3.5 Sonnet 91.6 — tie at top).

打开官方来源

Llama 3.1 70B

Llama 3.1 70B 为该系列中坚档:升级至 128K 上下文与多语言。亮点如 MMLU 86.0%、GSM8K 95.1%。

输入模态
文本
上下文
128K
参数
70B dense
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu 86 模型 llama-3-1-70b · 版本 0-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MMLU (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MMLU (0-shot, CoT) | 86

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 79.9 | GPT 3.5 Turbo 69.8。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

mmlu-pro 66.4 模型 llama-3-1-70b · 版本 5-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MMLU PRO (5-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MMLU PRO (5-shot, CoT) | 66.4

{
  "harness": null,
  "tools": null,
  "shots": 5,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 56.3 | GPT 3.5 Turbo 49.2。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

ifeval 87.5 模型 llama-3-1-70b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: IFEval · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: IFEval | 87.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 72.7 | GPT 3.5 Turbo 69.9。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

humaneval 80.5 模型 llama-3-1-70b · 版本 0-shot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: HumanEval (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: HumanEval (0-shot) | 80.5

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 75.6 | GPT 3.5 Turbo 68.0。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

mbpp-evalplus 86 模型 llama-3-1-70b · 版本 base, 0-shot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MBPP EvalPlus (base, 3-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MBPP EvalPlus (base, 3-shot) | 86

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 78.6 | GPT 3.5 Turbo 82.0。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

gsm8k 95.1 模型 llama-3-1-70b · 版本 8-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: GSM8K (8-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: GSM8K (8-shot, CoT) | 95.1

{
  "harness": null,
  "tools": null,
  "shots": 8,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 88.2 | GPT 3.5 Turbo 81.6。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

math 68 模型 llama-3-1-70b · 版本 0-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MATH (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MATH (0-shot, CoT) | 68

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 54.1 | GPT 3.5 Turbo 43.1。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

arc-challenge 94.8 模型 llama-3-1-70b · 版本 0-shot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: ARC Challenge (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: ARC Challenge (0-shot) | 94.8

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 88.7 | GPT 3.5 Turbo 83.7。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

gpqa 46.7 模型 llama-3-1-70b · 版本 0-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: GPQA (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: GPQA (0-shot, CoT) | 46.7

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 33.3 | GPT 3.5 Turbo 30.8。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

bfcl 84.8 模型 llama-3-1-70b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: BFCL · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: BFCL | 84.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct – | GPT 3.5 Turbo 85.9。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

nexus 56.7 模型 llama-3-1-70b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: Nexus · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: Nexus | 56.7

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 48.5 | GPT 3.5 Turbo 37.2。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

zeroscrolls-quality 90.5 模型 llama-3-1-70b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: ZeroSCROLLS/QuALITY · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: ZeroSCROLLS/QuALITY | 90.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct – | GPT 3.5 Turbo –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

infinitebench-en-mc 78.2 模型 llama-3-1-70b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: InfiniteBench/En.MC · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: InfiniteBench/En.MC | 78.2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct – | GPT 3.5 Turbo –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

niah 97.5 模型 llama-3-1-70b · 版本 NIH/Multi-needle · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: NIH/Multi-needle · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: NIH/Multi-needle | 97.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct – | GPT 3.5 Turbo –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

mgsm 86.9 模型 llama-3-1-70b · 版本 Multilingual, 0-shot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: Multilingual MGSM (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: Multilingual MGSM (0-shot) | 86.9

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Mixtral 8x22B Instruct 71.1 | GPT 3.5 Turbo 51.4。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

Llama 3.1 8B

Llama 3.1 8B 为该系列轻量档:升级至 128K 上下文与多语言。亮点如 GSM8K 84.5%、HumanEval 72.6%。

输入模态
文本
上下文
128K
参数
8B dense
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu 73 模型 llama-3-1-8b · 版本 0-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MMLU (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MMLU (0-shot, CoT) | 73

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 72.3(脚注: 5-shot, non-CoT) | Mistral 7B Instruct 60.5。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

mmlu-pro 48.3 模型 llama-3-1-8b · 版本 5-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MMLU PRO (5-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MMLU PRO (5-shot, CoT) | 48.3

{
  "harness": null,
  "tools": null,
  "shots": 5,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct 36.9。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

ifeval 80.4 模型 llama-3-1-8b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: IFEval · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: IFEval | 80.4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 73.6 | Mistral 7B Instruct 57.6。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

humaneval 72.6 模型 llama-3-1-8b · 版本 0-shot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: HumanEval (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: HumanEval (0-shot) | 72.6

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 54.3 | Mistral 7B Instruct 40.2。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

mbpp-evalplus 72.8 模型 llama-3-1-8b · 版本 base, 0-shot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MBPP EvalPlus (base, 3-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MBPP EvalPlus (base, 3-shot) | 72.8

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 71.7 | Mistral 7B Instruct 49.5。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

gsm8k 84.5 模型 llama-3-1-8b · 版本 8-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: GSM8K (8-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: GSM8K (8-shot, CoT) | 84.5

{
  "harness": null,
  "tools": null,
  "shots": 8,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 76.7 | Mistral 7B Instruct 53.2。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

math 51.9 模型 llama-3-1-8b · 版本 0-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: MATH (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: MATH (0-shot, CoT) | 51.9

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 44.3 | Mistral 7B Instruct 13.0。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

arc-challenge 83.4 模型 llama-3-1-8b · 版本 0-shot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: ARC Challenge (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: ARC Challenge (0-shot) | 83.4

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 87.6 | Mistral 7B Instruct 74.2。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

gpqa 32.8 模型 llama-3-1-8b · 版本 0-shot CoT · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: GPQA (0-shot, CoT) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: GPQA (0-shot, CoT) | 32.8

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": "CoT",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct 28.8。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

bfcl 76.1 模型 llama-3-1-8b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: BFCL · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: BFCL | 76.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct 60.4。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

nexus 38.5 模型 llama-3-1-8b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: Nexus · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: Nexus | 38.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 30.0 | Mistral 7B Instruct 24.7。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

zeroscrolls-quality 81 模型 llama-3-1-8b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: ZeroSCROLLS/QuALITY · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: ZeroSCROLLS/QuALITY | 81

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

infinitebench-en-mc 65.1 模型 llama-3-1-8b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: InfiniteBench/En.MC · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: InfiniteBench/En.MC | 65.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

niah 98.8 模型 llama-3-1-8b · 版本 NIH/Multi-needle · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: NIH/Multi-needle · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: NIH/Multi-needle | 98.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT – | Mistral 7B Instruct –。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源

mgsm 68.9 模型 llama-3-1-8b · 版本 Multilingual, 0-shot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Model evaluations · table: 8B/70B benchmark 总表(fbcdn PNG 渲染图) · row: Multilingual MGSM (0-shot) · figure: Model evaluations 区第二张 benchmark 总表图(live 页第 4 张 img, 3840x2040; 列: Llama 3.1 8B | Gemma 2 9B IT | Mistral 7B Instruct | Llama 3.1 70B | Mixtral 8x22B Instruct | GPT 3.5 Turbo) · quote_snippet: Multilingual MGSM (0-shot) | 68.9

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写自 live 页 8B/70B benchmark 总表(2026-09-01 audit 补行,70B 列与 batch-2 notes 转写逐格一致,8B 列为本轮独立读取)。对照列: Gemma 2 9B IT 53.2 | Mistral 7B Instruct 29.9。与 405B 同名行共享 benchmark_id/variant,protocol 一致。

打开官方来源