MiMo-7B-RL
Xiaomi / 小米 · 2025-04-30 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
MiMo-7B-RL
MiMo-7B-RL 是小米开源的 7B 强化学习对齐大模型,在 MMLU 达到 70.3 分,MATH-500 71.2 分,GSM8K 87.8 分,显著强化长思维链与代码数学逻辑能力。
- 输入模态
- 文本
- 上下文
- 32K
- 参数
- 7B
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Table 1: Main benchmark performance · row: MiMo-7B-RL
{
"harness": "lm-evaluation-harness",
"tools": null,
"shots": 5,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}5-shot 评估得分 70.3%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Table 1: Main benchmark performance · row: MiMo-7B-RL
{
"harness": "lm-evaluation-harness",
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}GPQA Diamond 子集 0-shot 准确率 38.5%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Table 1: Main benchmark performance · row: MiMo-7B-RL
{
"harness": "math-eval",
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}MATH-500 测试集 pass@1 得分 71.2%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Table 1: Main benchmark performance · row: MiMo-7B-RL
{
"harness": "aime-eval",
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}AIME 2024 官方实测 26.7%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Table 1: Main benchmark performance · row: MiMo-7B-RL
{
"harness": "livecodebench",
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}LiveCodeBench pass@1 准确率 32.4%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark · table: Reasoning Performance · row: MiMo-7B-RL
{
"harness": "lm-evaluation-harness",
"tools": null,
"shots": 8,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}GSM8K 8-shot 得分 87.8%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark · table: Coding Performance · row: MiMo-7B-RL
{
"harness": "evalplus",
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}HumanEval pass@1 得分 76.2%。