MiMo-V2.6-Pro / MiMo-V2.6-Flash / MiMo-V2.6-Pro-UltraSpeed
Xiaomi / 小米 · 2026-09-22 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
MiMo-V2.6-Pro
小米将 MiMo-V2.6-Pro 定位为该系列迄今能力最强的原生全模态模型,面向长程任务、网络安全与科研。发布页收录的评测以编码、通用智能体和安全为主,附录表 DeepSWE v1.1 为 71.9,Terminal-Bench 2.1 为 89.9。
- 输入模态
- 文本 / 图像 / 视频 / 音频
- 上下文
- 1M
- 参数
- 1.02T-A42B
- 价格(每百万 tokens)
- USD 输入 0.435 / 输出 0.87 · cache miss 输入价。cache hit 输入 $0.0036 / 百万 token。cache write 限时免费。发布页定价表。
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: DeepSWE v1.1 · quote_snippet: DeepSWE v1.1 | MiMo-V2.6-Pro 71.9
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 19.0;DeepSeek V4.1 Flash 74.2;Kimi K3 69.0;Claude Opus 5 74.0;Claude Fable 5 70.0;GPT 6 Astra 74.0。模型卡把 GPT-5.6 Sol 印为 73.0,发布页附录这一列为空,不并成一个数。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: ProgramBench · quote_snippet: ProgramBench | MiMo-V2.6-Pro 26.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 12.5;DeepSeek V4.1 Flash 20.3;Kimi K3 24.5;Claude Opus 5 37.0;GPT 5.6 Sol 25.0;Claude Fable 5 33.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: MiMo Code Bench · quote_snippet: MiMo Code Bench | MiMo-V2.6-Pro 63.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 40.4;DeepSeek V4.1 Flash 60.2;Kimi K3 60.1;Claude Opus 5 68.6;GPT 5.6 Sol 59.3;GPT 6 Astra 61.4。new-benchmark: mimo-code-bench。页面标注 in-house。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: GDPVal 2.1 (AA) · quote_snippet: GDPVal 2.1 (AA) | MiMo-V2.6-Pro 1673
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 1107;DeepSeek V4.1 Flash 1600;Kimi K3 1524;Claude Opus 5 1708;GPT 5.6 Sol 1588;Claude Fable 5 1595;GPT 6 Astra 1542;Claude Fable 5.1 1735。表注写明这是 Artificial Analysis 报告的 Elo。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: Toolathlon-verified · quote_snippet: Toolathlon-verified | MiMo-V2.6-Pro 76.9
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 49.1;Kimi K3 76.5;Claude Opus 5 80.6;GPT 5.6 Sol 74.9;Claude Fable 5 77.9;Claude Fable 5.1 77.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: Automation Bench v1.0.6 · quote_snippet: Automation Bench v1.0.6 | MiMo-V2.6-Pro 53.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 16.0;DeepSeek V4.1 Flash 54.8;Kimi K3 46.7;Claude Opus 5 50.3;GPT 5.6 Sol 45.8;Claude Fable 5 46.2;GPT 6 Astra 52.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: Agents' Last Exam · quote_snippet: Agents' Last Exam | MiMo-V2.6-Pro 31.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 13.2;DeepSeek V4.1 Flash 31.8;Kimi K3 28.3;Claude Opus 5 31.6;GPT 5.6 Sol 30.8;Claude Fable 5 25.7;GPT 6 Astra 34.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: Terminal Bench 4.0 · quote_snippet: Terminal Bench 4.0 | MiMo-V2.6-Pro 34.9
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 1.5;DeepSeek V4.1 Flash 26.8;Kimi K3 12.6;Claude Opus 5 49.0;GPT 5.6 Sol 39.9;Claude Fable 5 42.4;GPT 6 Astra 59.6;Claude Fable 5.1 55.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: Terminal Bench 2.1 · quote_snippet: Terminal Bench 2.1 | MiMo-V2.6-Pro 89.9
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 65.2;DeepSeek V4.1 Flash 90.6;Kimi K3 88.3;Claude Opus 5 89.1;GPT 5.6 Sol 88.8;Claude Fable 5 84.3;GPT 6 Astra 89.9;Claude Fable 5.1 91.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: OSWorld-Verified · quote_snippet: OSWorld-Verified | MiMo-V2.6-Pro 82.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:Kimi K3 84.8;Claude Opus 5 83.4;GPT 5.6 Sol 83.0;Claude Fable 5 86.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: JobBench · quote_snippet: JobBench | MiMo-V2.6-Pro 62.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 25.0;DeepSeek V4.1 Flash 45.8;Kimi K3 54.3;Claude Opus 5 65.7;GPT 5.6 Sol 45.4;Claude Fable 5 57.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: MiMo Visual Coding · quote_snippet: MiMo Visual Coding | MiMo-V2.6-Pro 72.3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:DeepSeek V4.1 Flash 70.6;Kimi K3 70.3;Claude Opus 5 70.0;GPT 5.6 Sol 73.4;Claude Fable 5 69.1;GPT 6 Astra 82.2;Claude Fable 5.1 74.4。new-benchmark: mimo-visual-coding。页面标注 in-house。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: CyberGym · quote_snippet: CyberGym | MiMo-V2.6-Pro 94.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 40.0;DeepSeek V4.1 Flash 88.1;Kimi K3 80.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: ExploitGym · quote_snippet: ExploitGym | MiMo-V2.6-Pro 17.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 0.1;DeepSeek V4.1 Flash 15.3;Kimi K3 8.1;Claude Opus 5 22.1;GPT 5.6 Sol 30.3;Claude Fable 5 28.4;GPT 6 Astra 42.4;Claude Fable 5.1 30.4。模型卡把 MiMo-V2.5-Pro 印为 0.2,发布页附录为 0.1。该列不是本系列分数。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: ExploitBench · quote_snippet: ExploitBench | MiMo-V2.6-Pro 47.9
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 16.6;Kimi K3 32.2;Claude Opus 5 70.0;GPT 5.6 Sol 78.5;Claude Fable 5 78.0;GPT 6 Astra 100.0;Claude Fable 5.1 83.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: SEC Bench Pro · quote_snippet: SEC Bench Pro | MiMo-V2.6-Pro 66.3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 17.7;DeepSeek V4.1 Flash 62.8;GPT 5.6 Sol 79.1;GPT 6 Astra 85.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: MiMo Cyber Bench · quote_snippet: MiMo Cyber Bench | MiMo-V2.6-Pro 81.7
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 0.0;DeepSeek V4.1 Flash 62.7;Kimi K3 56.3。new-benchmark: mimo-cyber-bench。页面标注 in-house。Pro 列与模型卡 80.2 不一致,模型卡另记一行。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Scaling RL, Fully Open-Sourced · row: DeepSWE v1.1 · quote_snippet: from 58.4 to 72.57
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}正文写 held-out DeepSWE v1.1 上 Pro 从 58.4 到 72.57。附录对比表同名行是 71.9,两处分开记。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pushing the Pareto Frontier · row: Artificial Analysis Intelligence Index · figure: images/aaindex-trim.png · quote_snippet: MiMo-V2.6-Pro scores 46.32 on the Artificial Analysis Intelligence Index
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}正文数字 46.32。柱状图替代文本写成 46,图只作示意,不另记一行。图注版本为 Intelligence Index v4.3,2026 年 9 月。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3. Evaluation Results · table: Evaluation Results · row: MiMo Cyber Bench · quote_snippet: MiMo Cyber Bench | 80.2 | 77.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mimo-cyber-bench。Pro-RL 与 Flash-RL 两张模型卡都印 Pro 80.2、Flash 77.2。发布页附录 Pro 为 81.7。Flash 两边同为 77.2,不重复记。
MiMo-V2.6-Flash
MiMo-V2.6-Flash 在发布文里被定位为智能、效率与成本更均衡的原生全模态模型。同页附录表 DeepSWE v1.1 为 67.9,Terminal-Bench 2.1 为 87.6。
- 输入模态
- 文本 / 图像 / 视频 / 音频
- 上下文
- 1M
- 参数
- 309B-A15B
- 价格(每百万 tokens)
- USD 输入 0.14 / 输出 0.28 · cache miss 输入价。cache hit 输入 $0.0028 / 百万 token。cache write 限时免费。发布页定价表。
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: DeepSWE v1.1 · quote_snippet: DeepSWE v1.1 | MiMo-V2.6-Flash 67.9
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 19.0;DeepSeek V4.1 Flash 74.2;Kimi K3 69.0;Claude Opus 5 74.0;Claude Fable 5 70.0;GPT 6 Astra 74.0。模型卡把 GPT-5.6 Sol 印为 73.0,发布页附录这一列为空,不并成一个数。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: ProgramBench · quote_snippet: ProgramBench | MiMo-V2.6-Flash 26.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 12.5;DeepSeek V4.1 Flash 20.3;Kimi K3 24.5;Claude Opus 5 37.0;GPT 5.6 Sol 25.0;Claude Fable 5 33.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: MiMo Code Bench · quote_snippet: MiMo Code Bench | MiMo-V2.6-Flash 61.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 40.4;DeepSeek V4.1 Flash 60.2;Kimi K3 60.1;Claude Opus 5 68.6;GPT 5.6 Sol 59.3;GPT 6 Astra 61.4。new-benchmark: mimo-code-bench。页面标注 in-house。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: Toolathlon-verified · quote_snippet: Toolathlon-verified | MiMo-V2.6-Flash 73.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 49.1;Kimi K3 76.5;Claude Opus 5 80.6;GPT 5.6 Sol 74.9;Claude Fable 5 77.9;Claude Fable 5.1 77.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: Automation Bench v1.0.6 · quote_snippet: Automation Bench v1.0.6 | MiMo-V2.6-Flash 52.3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 16.0;DeepSeek V4.1 Flash 54.8;Kimi K3 46.7;Claude Opus 5 50.3;GPT 5.6 Sol 45.8;Claude Fable 5 46.2;GPT 6 Astra 52.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: Agents' Last Exam · quote_snippet: Agents' Last Exam | MiMo-V2.6-Flash 27.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 13.2;DeepSeek V4.1 Flash 31.8;Kimi K3 28.3;Claude Opus 5 31.6;GPT 5.6 Sol 30.8;Claude Fable 5 25.7;GPT 6 Astra 34.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: Terminal Bench 4.0 · quote_snippet: Terminal Bench 4.0 | MiMo-V2.6-Flash 28.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 1.5;DeepSeek V4.1 Flash 26.8;Kimi K3 12.6;Claude Opus 5 49.0;GPT 5.6 Sol 39.9;Claude Fable 5 42.4;GPT 6 Astra 59.6;Claude Fable 5.1 55.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: Terminal Bench 2.1 · quote_snippet: Terminal Bench 2.1 | MiMo-V2.6-Flash 87.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 65.2;DeepSeek V4.1 Flash 90.6;Kimi K3 88.3;Claude Opus 5 89.1;GPT 5.6 Sol 88.8;Claude Fable 5 84.3;GPT 6 Astra 89.9;Claude Fable 5.1 91.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: OSWorld-Verified · quote_snippet: OSWorld-Verified | MiMo-V2.6-Flash 80.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:Kimi K3 84.8;Claude Opus 5 83.4;GPT 5.6 Sol 83.0;Claude Fable 5 86.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: JobBench · quote_snippet: JobBench | MiMo-V2.6-Flash 61.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 25.0;DeepSeek V4.1 Flash 45.8;Kimi K3 54.3;Claude Opus 5 65.7;GPT 5.6 Sol 45.4;Claude Fable 5 57.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: MiMo Visual Coding · quote_snippet: MiMo Visual Coding | MiMo-V2.6-Flash 71.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:DeepSeek V4.1 Flash 70.6;Kimi K3 70.3;Claude Opus 5 70.0;GPT 5.6 Sol 73.4;Claude Fable 5 69.1;GPT 6 Astra 82.2;Claude Fable 5.1 74.4。new-benchmark: mimo-visual-coding。页面标注 in-house。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: CyberGym · quote_snippet: CyberGym | MiMo-V2.6-Flash 95.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 40.0;DeepSeek V4.1 Flash 88.1;Kimi K3 80.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: ExploitGym · quote_snippet: ExploitGym | MiMo-V2.6-Flash 6.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 0.1;DeepSeek V4.1 Flash 15.3;Kimi K3 8.1;Claude Opus 5 22.1;GPT 5.6 Sol 30.3;Claude Fable 5 28.4;GPT 6 Astra 42.4;Claude Fable 5.1 30.4。模型卡把 MiMo-V2.5-Pro 印为 0.2,发布页附录为 0.1。该列不是本系列分数。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: ExploitBench · quote_snippet: ExploitBench | MiMo-V2.6-Flash 25.3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 16.6;Kimi K3 32.2;Claude Opus 5 70.0;GPT 5.6 Sol 78.5;Claude Fable 5 78.0;GPT 6 Astra 100.0;Claude Fable 5.1 83.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: SEC Bench Pro · quote_snippet: SEC Bench Pro | MiMo-V2.6-Flash 47.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 17.7;DeepSeek V4.1 Flash 62.8;GPT 5.6 Sol 79.1;GPT 6 Astra 85.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Appendix: Full Benchmark Results · table: bench.js 渲染的附录表 · row: MiMo Cyber Bench · quote_snippet: MiMo Cyber Bench | MiMo-V2.6-Flash 77.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}同表其他列:MiMo-V2.5-Pro 0.0;DeepSeek V4.1 Flash 62.7;Kimi K3 56.3。new-benchmark: mimo-cyber-bench。页面标注 in-house。Pro 列与模型卡 80.2 不一致,模型卡另记一行。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Scaling RL, Fully Open-Sourced · row: DeepSWE v1.1 · quote_snippet: from 48.8 to 65.68
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}正文写 held-out DeepSWE v1.1 上 Flash 从 48.8 到 65.68。附录对比表同名行是 67.9,两处分开记。
MiMo-V2.6-Pro-UltraSpeed
MiMo-V2.6-Pro-UltraSpeed 是 Pro 的高速输出模式,发布页写同等质量下输出速度最高约 20 倍,用于实时交互。本次发布没有给出该模式单独的评测分数。
- 输入模态
- 文本 / 图像 / 视频 / 音频
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 4.35 / 输出 8.7 · cache miss 输入价。cache hit 输入 $0.036 / 百万 token。发布页称同等质量下输出速度最高约 20 倍。
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。