← 模型目录

MAI-Code-1.1-Flash

Microsoft AI / MAI · 2026-08-11 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

MAI-Code-1.1-Flash

发布文自述定位是内置在 GitHub Copilot 与 VS Code 里的 "small, efficient, coding workhorse"(小而高效的编码工作马)。卡上收录的 5 项评测覆盖真实 Issue 修复、终端操作与视觉建站:SWE-bench Verified 72.6%、Terminal Bench 2.1 62.9%(与前代同在 GitHub Copilot 生产 VS Code harness 下自报,1.0 为 71.6%/51.7%),价格约为 1.0 的四分之一。

输入模态
文本 / 图像
上下文
256K
参数
138B 总参(5B 激活),稀疏 MoE
价格(每百万 tokens)
USD 输入 0.2 / 输出 1.2 · 缓存输入 $0.02/M;GitHub Copilot 列表价(model card 定价栏标 "To be finalized",以 GitHub 列表价为准);annual Copilot 订阅 0.25× premium request multiplier;较 MAI-Code-1-Flash($0.75/$0.075/$4.50)降价约 73%

本变体的评测证据

swebench 72.6 模型 mai-code-1-1-flash · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Quality and performance evaluations · table: MAI-Code-1.1-Flash / MAI-Code-1-Flash / Haiku 4.5 / GPT 5.4 mini(Pass rate + Tokens usage 四列) · row: SWE-Bench Verified · page: 5 · quote_snippet: SWE-Bench Verified 72.6 8.6K 71.6 10.8K 69.8 20.9K 69.2 9.4K

{
  "harness": "VS Code-based GitHub Copilot production harness",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

前代基线 MAI-Code-1-Flash = 71.6(10.8K tokens/task,本模型 8.6K);同表竞品(卡上明示同 harness 同设置):Haiku 4.5 69.8(20.9K)、GPT 5.4 mini 69.2(9.4K)。Benchmarking methodology:端到端计通过率,含仓库上下文/工具调用/验证,非精简 benchmark 环境;tokens usage = 每完成任务平均 token 数,可比通率下越短越好。发布文另称整体每任务 token 用量相对 1.0 −25%。PDF 文本层机读干净(pypdf 无列粘连),行值逐字核对。

打开官方来源

terminalbench 62.9 模型 mai-code-1-1-flash · 版本 2.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Quality and performance evaluations · table: MAI-Code-1.1-Flash / MAI-Code-1-Flash / Haiku 4.5 / GPT 5.4 mini(Pass rate + Tokens usage 四列) · row: Terminal Bench 2.1 · page: 5 · quote_snippet: Terminal Bench 2.1 62.9 17.0K 51.7 14.2K 49.4 25.5K 60.7 21.9K

{
  "harness": "VS Code-based GitHub Copilot production harness",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

前代基线 MAI-Code-1-Flash = 51.7(14.2K tokens/task,本模型 17.0K——通过率升但单任务 token 变多,卡上原样);同表竞品(同 harness 同设置):Haiku 4.5 49.4(25.5K)、GPT 5.4 mini 60.7(21.9K)。发布文相对论断并入本行:"a 22% improvement on Terminal-Bench 2.1 in GitHub Copilot CLI"(相对 1.0、CLI 场景,无 CLI 绝对值发布;注意卡上写 Copilot CLI 支持为 planned for a later rollout,两口径并存)。

打开官方来源

text2webapp 74.1 模型 mai-code-1-1-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Enabling vision capabilities for coding tasks · table: MAI-Code-1.1-Flash / Haiku 4.5 / GPT 5.4 mini(Pass rate + Tokens usage 三列,分组 WebApp development (internal)) · row: Text2WebApp · page: 5 · quote_snippet: Text2WebApp 74.1 17.1K 11.5 60.0K 58.3 36.4K

{
  "harness": "VS Code-based GitHub Copilot production harness",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: Text2WebApp——微软厂商内部视觉建站评测(文本提示生成 WebApp,卡上分组标 "(internal)"),公开侧无数据集与方法学定义,仅本 model card 一行,故不建实体、由兜底页承载;vendor-internal。同表竞品(同 harness):Haiku 4.5 11.5(60.0K tokens)、GPT 5.4 mini 58.3(36.4K);本模型 tokens 17.1K。PDF 文本层机读干净,行值逐字核对。

打开官方来源

screenshot2webapp 42.1 模型 mai-code-1-1-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Enabling vision capabilities for coding tasks · table: MAI-Code-1.1-Flash / Haiku 4.5 / GPT 5.4 mini(Pass rate + Tokens usage 三列,分组 WebApp development (internal)) · row: ScreenShot2WebApp · page: 5 · quote_snippet: ScreenShot2WebApp 42.1 10.5K 10.0 33.9K 39.3 26.6K

{
  "harness": "VS Code-based GitHub Copilot production harness",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ScreenShot2WebApp——微软厂商内部视觉建站评测(截图生成 WebApp,卡上分组标 "(internal)"),公开侧无数据集与方法学定义,仅本 model card 一行,故不建实体、由兜底页承载;vendor-internal。同表竞品(同 harness):Haiku 4.5 10.0(33.9K tokens)、GPT 5.4 mini 39.3(26.6K);本模型 tokens 10.5K。

打开官方来源

vision2web 11.5 模型 mai-code-1-1-flash · 版本 Level3 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-02 · 距快照 32 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Enabling vision capabilities for coding tasks · table: MAI-Code-1.1-Flash / Haiku 4.5 / GPT 5.4 mini(Pass rate + Tokens usage 三列,分组 WebApp development (internal)) · row: Vision2Web——Vision2Web Level3 · page: 5 · quote_snippet: Vision2Web Vision2Web Level3 11.5 15.1K 13.7 36.5K 10.1 3.7K

{
  "harness": "VS Code-based GitHub Copilot production harness",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

benchmark_id 三步判定第 1 步命中既有实体 vision2web(zai-org 开源,L3=全栈档;本行行名 "Vision2Web Level3" 与其层级结构一致,且 11.5% 是三行视觉评测中最低分,符合该实体 "层级越高越长程" 的画像),故按 variant=Level3 映射而非另铸 id。存疑留档:卡上把该行分组在 "WebApp development (internal)" 下,不能完全排除微软内部同名评测的可能,映射置信度中,建议人工复核;若复核为内部同名集,改铸新 id 并在本行留 revisions。同表竞品(同 harness):Haiku 4.5 13.7(36.5K tokens,高于本模型)、GPT 5.4 mini 10.1(3.7K);本模型 tokens 15.1K。

打开官方来源