← 模型目录

MiMo-7B-RL

Xiaomi / 小米 · 2025-04-30 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

MiMo-7B-RL

MiMo-7B-RL 是小米开源的 7B 强化学习对齐大模型,在 MMLU 达到 70.3 分,MATH-500 71.2 分,GSM8K 87.8 分,显著强化长思维链与代码数学逻辑能力。

输入模态
文本
上下文
32K
参数
7B
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu 70.3 模型 mimo-7b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Table 1: Main benchmark performance · row: MiMo-7B-RL

{
  "harness": "lm-evaluation-harness",
  "tools": null,
  "shots": 5,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

5-shot 评估得分 70.3%。

打开官方来源

gpqa 38.5 模型 mimo-7b · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Table 1: Main benchmark performance · row: MiMo-7B-RL

{
  "harness": "lm-evaluation-harness",
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

GPQA Diamond 子集 0-shot 准确率 38.5%。

打开官方来源

math500 71.2 模型 mimo-7b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Table 1: Main benchmark performance · row: MiMo-7B-RL

{
  "harness": "math-eval",
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

MATH-500 测试集 pass@1 得分 71.2%。

打开官方来源

aime24 26.7 模型 mimo-7b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Table 1: Main benchmark performance · row: MiMo-7B-RL

{
  "harness": "aime-eval",
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

AIME 2024 官方实测 26.7%。

打开官方来源

lcb 32.4 模型 mimo-7b · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Table 1: Main benchmark performance · row: MiMo-7B-RL

{
  "harness": "livecodebench",
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

LiveCodeBench pass@1 准确率 32.4%。

打开官方来源

gsm8k 87.8 模型 mimo-7b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark · table: Reasoning Performance · row: MiMo-7B-RL

{
  "harness": "lm-evaluation-harness",
  "tools": null,
  "shots": 8,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

GSM8K 8-shot 得分 87.8%。

打开官方来源

humaneval 76.2 模型 mimo-7b · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark · table: Coding Performance · row: MiMo-7B-RL

{
  "harness": "evalplus",
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

HumanEval pass@1 得分 76.2%。

打开官方来源