Gemma 4 E2B / Gemma 4 E4B / Gemma 4 26B MoE / Gemma 4 31B Dense
Google DeepMind / Gemini · 2026-04-02 · 通用模型
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Gemma 4 E2B
面向本地运行的开放权重模型,支持推理与工具调用;本变体输入能力和上下文按官方发布页记录。
- 输入模态
- 文本 / 图像 / 视频 / 音频
- 上下文
- 128K
- 参数
- Effective 2B
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMLU Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: AIME 2026 no tools
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: LiveCodeBench v6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: Codeforces ELO
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: Tau2 (average over 3)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: BigBench Extra Hard
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMMLU
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMMU Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: OmniDocBench 1.5 (average edit distance, lower is better)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MATH-Vision
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MedXPertQA MM
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: CoVoST
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。 官方仅标 CoVoST,未明示 CoVoST 2,不能静默归并到 covost2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: FLEURS (lower is better)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MRCR v2 8 needle 128k (average)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
Gemma 4 E4B
面向本地运行的开放权重模型,支持推理与工具调用;本变体输入能力和上下文按官方发布页记录。
- 输入模态
- 文本 / 图像 / 视频 / 音频
- 上下文
- 128K
- 参数
- Effective 4B
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMLU Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: AIME 2026 no tools
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: LiveCodeBench v6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: Codeforces ELO
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: Tau2 (average over 3)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: BigBench Extra Hard
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMMLU
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMMU Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: OmniDocBench 1.5 (average edit distance, lower is better)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MATH-Vision
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MedXPertQA MM
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: CoVoST
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。 官方仅标 CoVoST,未明示 CoVoST 2,不能静默归并到 covost2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: FLEURS (lower is better)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MRCR v2 8 needle 128k (average)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
Gemma 4 26B MoE
面向本地运行的开放权重模型,支持推理与工具调用;本变体输入能力和上下文按官方发布页记录。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 256K
- 参数
- 26B MoE; active 3.8B
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMLU Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: AIME 2026 no tools
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: LiveCodeBench v6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: Codeforces ELO
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: Tau2 (average over 3)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: HLE no tools
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: HLE with search
{
"harness": null,
"tools": [
"search"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: BigBench Extra Hard
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMMLU
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMMU Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: OmniDocBench 1.5 (average edit distance, lower is better)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MATH-Vision
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MedXPertQA MM
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MRCR v2 8 needle 128k (average)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
Gemma 4 31B Dense
面向本地运行的开放权重模型,支持推理与工具调用;本变体输入能力和上下文按官方发布页记录。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 256K
- 参数
- 31B Dense
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMLU Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: AIME 2026 no tools
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: LiveCodeBench v6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: Codeforces ELO
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: Tau2 (average over 3)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: HLE no tools
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: HLE with search
{
"harness": null,
"tools": [
"search"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: BigBench Extra Hard
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 new-benchmark: 官方页面使用的评测名称。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMMLU
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MMMU Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: OmniDocBench 1.5 (average edit distance, lower is better)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MATH-Vision
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MedXPertQA MM
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Benchmark Results · row: MRCR v2 8 needle 128k (average)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}协议仅记录官方明示字段,其余留空;不作跨协议排名。 取 2026-09-23 模型卡快照;不代表发布日原始表。