← 模型目录

MAI-Transcribe-2-Streaming

Microsoft AI / MAI · 2026-10-01 · 专项模型

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

MAI-Transcribe-2-Streaming

微软 AI 的首款流式语音识别模型,支持 60 种语言的低延迟实时转写与自动连续语种检测,约 100ms 给出首批预测假设,Artificial Analysis 流式最终词错误率 2.50%、首部词错误率 2.80% 均位列第一,最终时间延迟 0.130s 位于帕累托前沿。

输入模态
音频 / 文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 未说明 / 输出 未说明 · 首发介绍期每小时音频 $0.54(截至年底)

本变体的评测证据

artificial-analysis-wer-streaming 2.50% 模型 mai-transcribe-2-streaming · 版本 Final WER, streaming · 指标 lower is better · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-10-04 · 距快照 0 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Meet MAI-Transcribe-2-Streaming: Real-time transcriptions. Really fast. · row: Final WER, streaming · quote_snippet: ranks no. 1 for accuracy for both final and partial transcripts on Artificial Analysis

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: Artificial Analysis 流式语音识别基准;最终词错误率(Final WER)2.50%,排名第一;Grok Voice Transcribe 2.0 为 2.73%,Muse Voice Transcribe 为 3.06%,Cartesia Ink Preview 为 3.11%。

打开官方来源

artificial-analysis-wer-streaming 2.80% 模型 mai-transcribe-2-streaming · 版本 First Partial WER, streaming · 指标 lower is better · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-10-04 · 距快照 0 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Meet MAI-Transcribe-2-Streaming: Real-time transcriptions. Really fast. · row: First Partial WER, streaming · quote_snippet: ranks no. 1 for accuracy for both final and partial transcripts on Artificial Analysis

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: 首部预测词错误率(First Partial WER)2.80%,排名第一;Grok 2.0 为 3.36%,Muse 为 3.57%,Cartesia Ink-2 为 4.89%。

打开官方来源

artificial-analysis-time-to-final 0.130s 模型 mai-transcribe-2-streaming · 版本 Time to Final, streaming · 指标 lower is better · 单位 seconds 来源等级 A · third_party_reported · 核对日期 2026-10-04 · 距快照 0 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Meet MAI-Transcribe-2-Streaming: Real-time transcriptions. Really fast. · row: Time to Final, streaming · quote_snippet: And on its accuracy-versus-latency evaluation, we sit on the Pareto frontier

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: 达到最终转写所需时间为 0.130 秒(130ms),位于帕累托前沿。

打开官方来源

turing-test-speech 50.3% 模型 mai-transcribe-2-streaming · 版本 4000-listener Turing Test · 指标 higher is better · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-04 · 距快照 0 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: MAI-Voice-2.1: Seamlessly support multilingual experiences · row: 4000-listener Turing Test · quote_snippet: In a 4,000-listener Turing Test, 50.3% rated MAI-Voice as equally or more human-like than human recordings

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: 4000 名听众盲测图灵测试,50.3% 认为语音与真人录音同样或更像真人;统计综合了 MAI-Voice-2.1 与 MAI-Voice-2.1-Flash。

打开官方来源