MiMo-V2.5-Pro
Xiaomi / 小米 · 2026-03-19 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
MiMo-V2.5-Pro
MiMo-V2.5-Pro 是小米万亿级参数旗舰混合专家大模型,总参数 1.02T,激活 42B,支持 100 万 token 上下文。评测全面对标国际前沿:MMLU 89.4、GSM8K 99.6、MATH 86.2、SWE-bench Verified 78.9%、C-Eval 91.5、BBH 88.4。
- 输入模态
- 文本 / 代码 / 多模态
- 上下文
- 1M
- 参数
- 1.02T-A42B
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro
{
"harness": "lm-evaluation-harness",
"tools": null,
"shots": 5,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}MMLU 5-shot 官方实测得分 89.4%(MMLU-Redux 92.8%)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro
{
"harness": "mmlu-pro",
"tools": null,
"shots": 5,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}MMLU-Pro 5-shot 官方得分 68.5%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 2: Mathematical Reasoning · row: MiMo-V2.5-Pro
{
"harness": "lm-evaluation-harness",
"tools": null,
"shots": 8,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}GSM8K 8-shot 得分 99.6%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 2: Mathematical Reasoning · row: MiMo-V2.5-Pro
{
"harness": "math-eval",
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}MATH 全集 0-shot 得分 86.2%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro
{
"harness": "lm-evaluation-harness",
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}GPQA Diamond 0-shot 准确率 66.7%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 2: Mathematical Reasoning · row: MiMo-V2.5-Pro
{
"harness": "aime-eval",
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}AIME 24&25 官方实测得分 37.3%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 3: Code Generation and Reasoning · row: MiMo-V2.5-Pro
{
"harness": "evalplus",
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}HumanEval+ pass@1 准确率 75.6%(MBPP+ 74.1%)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 3: Code Generation and Reasoning · row: MiMo-V2.5-Pro
{
"harness": "livecodebench-v6",
"tools": null,
"shots": 0,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}LiveCodeBench v6 pass@1 准确率 39.6%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 3: Code Generation and Reasoning · row: MiMo-V2.5-Pro
{
"harness": "agentless-framework",
"tools": "bash,python",
"shots": null,
"reasoning_effort": null,
"temperature": 0.2,
"top_p": 0.95,
"token_budget": null,
"turn_limit": 30,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": "pytest"
}SWE-bench Verified 解决率 78.9%(AgentLess 下为 35.7%)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 4: Chinese Language Understanding · row: MiMo-V2.5-Pro
{
"harness": "ceval-eval",
"tools": null,
"shots": 5,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}C-Eval 5-shot 综合准确率 91.5%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 4: Chinese Language Understanding · row: MiMo-V2.5-Pro
{
"harness": "cmmlu-eval",
"tools": null,
"shots": 5,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}CMMLU 5-shot 得分 90.2%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro
{
"harness": "lm-evaluation-harness",
"tools": null,
"shots": 3,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}Big-Bench Hard (BBH) 3-shot 准确率 88.4%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro
{
"harness": "lm-evaluation-harness",
"tools": null,
"shots": 3,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}DROP 3-shot F1 得分 86.3%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro
{
"harness": "lm-evaluation-harness",
"tools": null,
"shots": 25,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}ARC-Challenge 25-shot 得分 97.2%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark Results · table: Table 1: Language and Knowledge Performance · row: MiMo-V2.5-Pro
{
"harness": "lm-evaluation-harness",
"tools": null,
"shots": 10,
"reasoning_effort": null,
"temperature": 0,
"top_p": 1,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 1,
"aggregation": "mean",
"judge": null
}HellaSwag 10-shot 得分 89.8%。