← 模型目录

MiMo-V2.5-Pro

Xiaomi / 小米 · 2026-03-19 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

MiMo-V2.5-Pro

MiMo-V2.5-Pro 是小米万亿级参数旗舰混合专家大模型,总参数 1.02T,激活 42B,支持 100 万 token 上下文。评测全面对标国际前沿:MMLU 89.4、GSM8K 99.6、MATH 86.2、SWE-bench Verified 78.9%、C-Eval 91.5、BBH 88.4。

输入模态
文本 / 代码 / 多模态
上下文
1M
参数
1.02T-A42B
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu 89.4 模型 mimo-v2-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro

{
  "harness": "lm-evaluation-harness",
  "tools": null,
  "shots": 5,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

MMLU 5-shot 官方实测得分 89.4%(MMLU-Redux 92.8%)。

打开官方来源

mmlu-pro 68.5 模型 mimo-v2-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro

{
  "harness": "mmlu-pro",
  "tools": null,
  "shots": 5,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

MMLU-Pro 5-shot 官方得分 68.5%。

打开官方来源

gsm8k 99.6 模型 mimo-v2-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 2: Mathematical Reasoning · row: MiMo-V2.5-Pro

{
  "harness": "lm-evaluation-harness",
  "tools": null,
  "shots": 8,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

GSM8K 8-shot 得分 99.6%。

打开官方来源

math 86.2 模型 mimo-v2-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 2: Mathematical Reasoning · row: MiMo-V2.5-Pro

{
  "harness": "math-eval",
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

MATH 全集 0-shot 得分 86.2%。

打开官方来源

gpqa 66.7 模型 mimo-v2-5-pro · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro

{
  "harness": "lm-evaluation-harness",
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

GPQA Diamond 0-shot 准确率 66.7%。

打开官方来源

aime24 37.3 模型 mimo-v2-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 2: Mathematical Reasoning · row: MiMo-V2.5-Pro

{
  "harness": "aime-eval",
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

AIME 24&25 官方实测得分 37.3%。

打开官方来源

humaneval 75.6 模型 mimo-v2-5-pro · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 3: Code Generation and Reasoning · row: MiMo-V2.5-Pro

{
  "harness": "evalplus",
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

HumanEval+ pass@1 准确率 75.6%(MBPP+ 74.1%)。

打开官方来源

lcb 39.6 模型 mimo-v2-5-pro · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 3: Code Generation and Reasoning · row: MiMo-V2.5-Pro

{
  "harness": "livecodebench-v6",
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

LiveCodeBench v6 pass@1 准确率 39.6%。

打开官方来源

swebench 78.9 模型 mimo-v2-5-pro · 版本 Verified · 指标 resolve_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 3: Code Generation and Reasoning · row: MiMo-V2.5-Pro

{
  "harness": "agentless-framework",
  "tools": "bash,python",
  "shots": null,
  "reasoning_effort": null,
  "temperature": 0.2,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": 30,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": "pytest"
}

SWE-bench Verified 解决率 78.9%(AgentLess 下为 35.7%)。

打开官方来源

ceval 91.5 模型 mimo-v2-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 4: Chinese Language Understanding · row: MiMo-V2.5-Pro

{
  "harness": "ceval-eval",
  "tools": null,
  "shots": 5,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

C-Eval 5-shot 综合准确率 91.5%。

打开官方来源

cmmlu 90.2 模型 mimo-v2-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 4: Chinese Language Understanding · row: MiMo-V2.5-Pro

{
  "harness": "cmmlu-eval",
  "tools": null,
  "shots": 5,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

CMMLU 5-shot 得分 90.2%。

打开官方来源

bbh 88.4 模型 mimo-v2-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro

{
  "harness": "lm-evaluation-harness",
  "tools": null,
  "shots": 3,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

Big-Bench Hard (BBH) 3-shot 准确率 88.4%。

打开官方来源

drop 86.3 模型 mimo-v2-5-pro · 版本 未说明 · 指标 F1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro

{
  "harness": "lm-evaluation-harness",
  "tools": null,
  "shots": 3,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

DROP 3-shot F1 得分 86.3%。

打开官方来源

arc-challenge 97.2 模型 mimo-v2-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro

{
  "harness": "lm-evaluation-harness",
  "tools": null,
  "shots": 25,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

ARC-Challenge 25-shot 得分 97.2%。

打开官方来源

hellaswag 89.8 模型 mimo-v2-5-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-12 · 距快照 22 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro

{
  "harness": "lm-evaluation-harness",
  "tools": null,
  "shots": 10,
  "reasoning_effort": null,
  "temperature": 0,
  "top_p": 1,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "mean",
  "judge": null
}

HellaSwag 10-shot 得分 89.8%。

打开官方来源