← 模型目录

Gemma 4 E2B / Gemma 4 E4B / Gemma 4 26B MoE / Gemma 4 31B Dense

Google DeepMind / Gemini · 2026-04-02 · 通用模型

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Gemma 4 E2B

面向本地运行的开放权重模型,支持推理与工具调用;本变体输入能力和上下文按官方发布页记录。

输入模态
文本 / 图像 / 视频 / 音频
上下文
128K
参数
Effective 2B
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu-pro 60.0% 模型 gemma-4-e2b · 版本 MMLU Pro · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMLU Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

aime-26 37.5% 模型 gemma-4-e2b · 版本 AIME 2026 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: AIME 2026 no tools

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

lcb 44.0% 模型 gemma-4-e2b · 版本 LiveCodeBench v6 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: LiveCodeBench v6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

codeforces 633 模型 gemma-4-e2b · 版本 Codeforces ELO · 指标 elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: Codeforces ELO

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

gpqa 43.4% 模型 gemma-4-e2b · 版本 GPQA Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

tau-bench 24.5% 模型 gemma-4-e2b · 版本 Tau2 (average over 3) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: Tau2 (average over 3)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

bbeh 21.9% 模型 gemma-4-e2b · 版本 BigBench Extra Hard · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: BigBench Extra Hard

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mmmlu 67.4% 模型 gemma-4-e2b · 版本 MMMLU · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMMLU

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mmmu 44.2% 模型 gemma-4-e2b · 版本 MMMU Pro · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMMU Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

omnidocbench 0.290 模型 gemma-4-e2b · 版本 OmniDocBench 1.5 (average edit distance, lower is better) · 指标 average_edit_distance · 单位 未说明 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: OmniDocBench 1.5 (average edit distance, lower is better)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mathvision 52.4% 模型 gemma-4-e2b · 版本 MATH-Vision · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MATH-Vision

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

medxpertqa-mm 23.5% 模型 gemma-4-e2b · 版本 MedXPertQA MM · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MedXPertQA MM

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

covost 33.47 模型 gemma-4-e2b · 版本 CoVoST · 指标 未说明 · 单位 未说明 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: CoVoST

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。 官方仅标 CoVoST,未明示 CoVoST 2,不能静默归并到 covost2。

打开官方来源

fleurs 0.09 模型 gemma-4-e2b · 版本 FLEURS (lower is better) · 指标 未说明 · 单位 未说明 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: FLEURS (lower is better)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mrcr 19.1% 模型 gemma-4-e2b · 版本 MRCR v2 8 needle 128k (average) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MRCR v2 8 needle 128k (average)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

Gemma 4 E4B

面向本地运行的开放权重模型,支持推理与工具调用;本变体输入能力和上下文按官方发布页记录。

输入模态
文本 / 图像 / 视频 / 音频
上下文
128K
参数
Effective 4B
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu-pro 69.4% 模型 gemma-4-e4b · 版本 MMLU Pro · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMLU Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

aime-26 42.5% 模型 gemma-4-e4b · 版本 AIME 2026 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: AIME 2026 no tools

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

lcb 52.0% 模型 gemma-4-e4b · 版本 LiveCodeBench v6 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: LiveCodeBench v6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

codeforces 940 模型 gemma-4-e4b · 版本 Codeforces ELO · 指标 elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: Codeforces ELO

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

gpqa 58.6% 模型 gemma-4-e4b · 版本 GPQA Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

tau-bench 42.2% 模型 gemma-4-e4b · 版本 Tau2 (average over 3) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: Tau2 (average over 3)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

bbeh 33.1% 模型 gemma-4-e4b · 版本 BigBench Extra Hard · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: BigBench Extra Hard

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mmmlu 76.6% 模型 gemma-4-e4b · 版本 MMMLU · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMMLU

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mmmu 52.6% 模型 gemma-4-e4b · 版本 MMMU Pro · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMMU Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

omnidocbench 0.181 模型 gemma-4-e4b · 版本 OmniDocBench 1.5 (average edit distance, lower is better) · 指标 average_edit_distance · 单位 未说明 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: OmniDocBench 1.5 (average edit distance, lower is better)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mathvision 59.5% 模型 gemma-4-e4b · 版本 MATH-Vision · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MATH-Vision

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

medxpertqa-mm 28.7% 模型 gemma-4-e4b · 版本 MedXPertQA MM · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MedXPertQA MM

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

covost 35.54 模型 gemma-4-e4b · 版本 CoVoST · 指标 未说明 · 单位 未说明 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: CoVoST

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。 官方仅标 CoVoST,未明示 CoVoST 2,不能静默归并到 covost2。

打开官方来源

fleurs 0.08 模型 gemma-4-e4b · 版本 FLEURS (lower is better) · 指标 未说明 · 单位 未说明 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: FLEURS (lower is better)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mrcr 25.4% 模型 gemma-4-e4b · 版本 MRCR v2 8 needle 128k (average) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MRCR v2 8 needle 128k (average)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

Gemma 4 26B MoE

面向本地运行的开放权重模型,支持推理与工具调用;本变体输入能力和上下文按官方发布页记录。

输入模态
文本 / 图像 / 视频
上下文
256K
参数
26B MoE; active 3.8B
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu-pro 82.6% 模型 gemma-4-26b · 版本 MMLU Pro · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMLU Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

aime-26 88.3% 模型 gemma-4-26b · 版本 AIME 2026 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: AIME 2026 no tools

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

lcb 77.1% 模型 gemma-4-26b · 版本 LiveCodeBench v6 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: LiveCodeBench v6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

codeforces 1718 模型 gemma-4-26b · 版本 Codeforces ELO · 指标 elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: Codeforces ELO

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

gpqa 82.3% 模型 gemma-4-26b · 版本 GPQA Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

tau-bench 68.2% 模型 gemma-4-26b · 版本 Tau2 (average over 3) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: Tau2 (average over 3)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

hlehle 8.7% 模型 gemma-4-26b · 版本 HLE no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: HLE no tools

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

bbeh 64.8% 模型 gemma-4-26b · 版本 BigBench Extra Hard · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: BigBench Extra Hard

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mmmlu 86.3% 模型 gemma-4-26b · 版本 MMMLU · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMMLU

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mmmu 73.8% 模型 gemma-4-26b · 版本 MMMU Pro · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMMU Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

omnidocbench 0.149 模型 gemma-4-26b · 版本 OmniDocBench 1.5 (average edit distance, lower is better) · 指标 average_edit_distance · 单位 未说明 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: OmniDocBench 1.5 (average edit distance, lower is better)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mathvision 82.4% 模型 gemma-4-26b · 版本 MATH-Vision · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MATH-Vision

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

medxpertqa-mm 58.1% 模型 gemma-4-26b · 版本 MedXPertQA MM · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MedXPertQA MM

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mrcr 44.1% 模型 gemma-4-26b · 版本 MRCR v2 8 needle 128k (average) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MRCR v2 8 needle 128k (average)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

Gemma 4 31B Dense

面向本地运行的开放权重模型,支持推理与工具调用;本变体输入能力和上下文按官方发布页记录。

输入模态
文本 / 图像 / 视频
上下文
256K
参数
31B Dense
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu-pro 85.2% 模型 gemma-4-31b · 版本 MMLU Pro · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMLU Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

aime-26 89.2% 模型 gemma-4-31b · 版本 AIME 2026 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: AIME 2026 no tools

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

lcb 80.0% 模型 gemma-4-31b · 版本 LiveCodeBench v6 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: LiveCodeBench v6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

codeforces 2150 模型 gemma-4-31b · 版本 Codeforces ELO · 指标 elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: Codeforces ELO

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

gpqa 84.3% 模型 gemma-4-31b · 版本 GPQA Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

tau-bench 76.9% 模型 gemma-4-31b · 版本 Tau2 (average over 3) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: Tau2 (average over 3)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

hlehle 19.5% 模型 gemma-4-31b · 版本 HLE no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: HLE no tools

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

bbeh 74.4% 模型 gemma-4-31b · 版本 BigBench Extra Hard · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: BigBench Extra Hard

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mmmlu 88.4% 模型 gemma-4-31b · 版本 MMMLU · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMMLU

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mmmu 76.9% 模型 gemma-4-31b · 版本 MMMU Pro · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MMMU Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

omnidocbench 0.131 模型 gemma-4-31b · 版本 OmniDocBench 1.5 (average edit distance, lower is better) · 指标 average_edit_distance · 单位 未说明 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: OmniDocBench 1.5 (average edit distance, lower is better)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mathvision 85.6% 模型 gemma-4-31b · 版本 MATH-Vision · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MATH-Vision

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

medxpertqa-mm 61.3% 模型 gemma-4-31b · 版本 MedXPertQA MM · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MedXPertQA MM

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源

mrcr 66.4% 模型 gemma-4-31b · 版本 MRCR v2 8 needle 128k (average) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-23 · 距快照 11 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Benchmark Results · row: MRCR v2 8 needle 128k (average)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。

打开官方来源