← 模型目录

MAI-Thinking-1

Microsoft AI / MAI · 2026-08-12 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

MAI-Thinking-1

Microsoft AI 将 MAI-Thinking-1 定位为面向复杂企业任务的推理模型——以成本高效的推理覆盖密集企业任务,在同级(weight class)数学、知识与编码上达到 SOTA。本次发布可核验的评测以数学与智能体编码为主:AIME 2025 97.0%、AIME 2026 94.5%(散文明文),SWE-Bench Pro 未报数值,散文以 toe-to-toe with Claude Opus 4.6 定位。

输入模态
文本
上下文
256K
参数
~1T 总参(35B 激活),稀疏 MoE
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

aime-25 97.0% 模型 mai-thinking-1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Advanced mathematical reasoning capabilities · quote_snippet: MAI-Thinking-1 reaches 97.0% on AIME 2025, and 94.5% on AIME 2026

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose plain-text score. The Table 1 figure carries the same benchmark with competitor columns (Sonnet 4.6 95.6 / Opus 4.6 99.8 / DeepSeek V3.2 93.1, visual read) — recorded as the separate pending row microsoft-mai-thinking-1--aime-25-table. [Table 1 视觉转写归并] 视觉转写(未人工确认):MAI-Thinking-1 97.0;竞品列 Sonnet 4.6 95.6 / Opus 4.6 99.8 / GPT 5.4 - / Kimi K2.6 - / DeepSeek V3.2 93.1 / DeepSeek V4 - / GLM-5.1 -。与散文 verified 行(97.0)一致。图内协议脚注(视觉转写):其余(非 agentic coding)评测使用 max output tokens 256k——未经人工确认,protocol 保持 null。升级路径:人工读图核对后翻转。

打开官方来源

aime-26 94.5% 模型 mai-thinking-1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Advanced mathematical reasoning capabilities · quote_snippet: and 94.5% on AIME 2026, showing strong mathematical and scientific reasoning for its weight class

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose plain-text score; benchmark entity aime-26 already exists in data/benchmarks/ (no new cast needed). Table 1 figure row with competitor columns (Kimi K2.6 96.4 / GLM-5.1 95.3, visual read) is the separate pending row microsoft-mai-thinking-1--aime-26-table. [Table 1 视觉转写归并] 视觉转写(未人工确认):MAI-Thinking-1 94.5;竞品列 Sonnet 4.6 - / Opus 4.6 - / GPT 5.4 - / Kimi K2.6 96.4 / DeepSeek V3.2 - / DeepSeek V4 - / GLM-5.1 95.3。与散文 verified 行(94.5)一致。升级路径:人工读图核对后翻转。

打开官方来源

swebench-pro 官方未公布数值 模型 mai-thinking-1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Medium-sized model, with strong software engineering performance · quote_snippet: our model is toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose states no number — only the positioning claim toe-to-toe with Claude Opus 4.6 (verified as a locatable claim, Opus 5 precedent). Hero subtitle separately claims top SWE-Bench Pro results at a mid-weight price. The Table 1 figure row (separate pending entry microsoft-mai-thinking-1--swebench-pro-table) visually reads 52.8 vs Opus 4.6 53.4, consistent with the toe-to-toe wording pending human read. [Table 1 视觉转写归并] 视觉转写(未人工确认):MAI-Thinking-1 52.8;竞品列 Sonnet 4.6 - / Opus 4.6 53.4 / GPT 5.4 57.7 / Kimi K2.6 58.6 / DeepSeek V3.2 - / DeepSeek V4 55.4 / GLM-5.1 58.4。跨厂锚点:GPT 5.4 57.7 与 OpenAI 自报 SWE-Bench Pro 一致。与散文 verified 行(toe-to-toe with Opus 4.6,无数字)同 benchmark 分立:text 行记论断、figure 行记数值,id 用 -table 后缀区分(kimi-k3 deepswe footnote/table 先例)。升级路径:人工读图核对后翻转。

打开官方来源

hmmt-26 图表尚无可读数值 模型 mai-thinking-1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Results — Table 1. MAI-Thinking-1 metrics · row: HMMT Feb 2026 (STEM section) · figure: MAI-Thinking-1-metrics.png — HMMT Feb 2026 行, MAI-Thinking-1 列 · quote_snippet: Table 1. MAI-Thinking-1 metrics — post-trained model evaluation results on public STEM and agentic coding benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写(未人工确认):MAI-Thinking-1 84.9;竞品列 Sonnet 4.6 - / Opus 4.6 - / GPT 5.4 - / Kimi K2.6 92.7 / DeepSeek V3.2 - / DeepSeek V4 95.2 / GLM-5.1 82.6。benchmark 实体 hmmt-26(HMMT February 2026)已存在,行名 HMMT Feb 2026 直接映射,无需新铸。升级路径:人工读图核对后翻转。

打开官方来源

gpqa 图表尚无可读数值 模型 mai-thinking-1 · 版本 Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Results — Table 1. MAI-Thinking-1 metrics · row: GPQA Diamond (STEM section) · figure: MAI-Thinking-1-metrics.png — GPQA Diamond 行, MAI-Thinking-1 列 · quote_snippet: Table 1. MAI-Thinking-1 metrics — post-trained model evaluation results on public STEM and agentic coding benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写(未人工确认):MAI-Thinking-1 84.2;竞品列 Sonnet 4.6 89.9 / Opus 4.6 91.3 / GPT 5.4 92.8 / Kimi K2.6 90.5 / DeepSeek V3.2 82.4 / DeepSeek V4 90.1 / GLM-5.1 86.2——本行 8 家全有值且 MAI-Thinking-1 垫底,为全表最需人工复核的一行。升级路径:人工读图核对后翻转。

打开官方来源

lcb 图表尚无可读数值 模型 mai-thinking-1 · 版本 v6 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Results — Table 1. MAI-Thinking-1 metrics · row: LCB v6 (STEM section) · figure: MAI-Thinking-1-metrics.png — LCB v6 行, MAI-Thinking-1 列 · quote_snippet: Table 1. MAI-Thinking-1 metrics — post-trained model evaluation results on public STEM and agentic coding benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写(未人工确认):MAI-Thinking-1 87.7;竞品列 Sonnet 4.6 - / Opus 4.6 - / GPT 5.4 - / Kimi K2.6 89.6 / DeepSeek V3.2 83.3 / DeepSeek V4 93.5 / GLM-5.1 -。LiveCodeBench v6 窗口记 variant。升级路径:人工读图核对后翻转。

打开官方来源

terminalbench 图表尚无可读数值 模型 mai-thinking-1 · 版本 2.0 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Results — Table 1. MAI-Thinking-1 metrics · row: Terminal Bench 2.0 (Agentic Coding section) · figure: MAI-Thinking-1-metrics.png — Terminal Bench 2.0 行, MAI-Thinking-1 列 · quote_snippet: Table 1. MAI-Thinking-1 metrics — post-trained model evaluation results on public STEM and agentic coding benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写(未人工确认):MAI-Thinking-1 46.0;竞品列 Sonnet 4.6 59.1 / Opus 4.6 65.4 / GPT 5.4 75.1 / Kimi K2.6 66.7 / DeepSeek V3.2 46.4 / DeepSeek V4 67.9 / GLM-5.1 69.0。跨厂锚点:GPT 5.4 75.1 与 OpenAI GPT-5.5 页 Terminal-Bench 列读数一致(彼处口径为 2.1);本表自述竞品数取自各家官方 model card,Terminal-Bench 版本口径可能不齐,跨厂比较须按 variant 区分。图内协议脚注(视觉转写):agentic coding 评测总上下文长度 256k——未经人工确认,protocol 保持 null。升级路径:人工读图核对后翻转。

打开官方来源

swebench 图表尚无可读数值 模型 mai-thinking-1 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Results — Table 1. MAI-Thinking-1 metrics · row: SWE-Bench Verified (Agentic Coding section) · figure: MAI-Thinking-1-metrics.png — SWE-Bench Verified 行, MAI-Thinking-1 列 · quote_snippet: Table 1. MAI-Thinking-1 metrics — post-trained model evaluation results on public STEM and agentic coding benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

视觉转写(未人工确认):MAI-Thinking-1 73.5;竞品列 Sonnet 4.6 79.6 / Opus 4.6 80.8 / GPT 5.4 - / Kimi K2.6 80.2 / DeepSeek V3.2 73.1 / DeepSeek V4 80.6 / GLM-5.1 -。升级路径:人工读图核对后翻转。

打开官方来源