MAI-Transcribe-2-Streaming
Microsoft AI / MAI · 2026-10-01 · 专项模型
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
MAI-Transcribe-2-Streaming
微软 AI 的首款流式语音识别模型,支持 60 种语言的低延迟实时转写与自动连续语种检测,约 100ms 给出首批预测假设,Artificial Analysis 流式最终词错误率 2.50%、首部词错误率 2.80% 均位列第一,最终时间延迟 0.130s 位于帕累托前沿。
- 输入模态
- 音频 / 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 未说明 / 输出 未说明 · 首发介绍期每小时音频 $0.54(截至年底)
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Meet MAI-Transcribe-2-Streaming: Real-time transcriptions. Really fast. · row: Final WER, streaming · quote_snippet: ranks no. 1 for accuracy for both final and partial transcripts on Artificial Analysis
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: Artificial Analysis 流式语音识别基准;最终词错误率(Final WER)2.50%,排名第一;Grok Voice Transcribe 2.0 为 2.73%,Muse Voice Transcribe 为 3.06%,Cartesia Ink Preview 为 3.11%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Meet MAI-Transcribe-2-Streaming: Real-time transcriptions. Really fast. · row: First Partial WER, streaming · quote_snippet: ranks no. 1 for accuracy for both final and partial transcripts on Artificial Analysis
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: 首部预测词错误率(First Partial WER)2.80%,排名第一;Grok 2.0 为 3.36%,Muse 为 3.57%,Cartesia Ink-2 为 4.89%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Meet MAI-Transcribe-2-Streaming: Real-time transcriptions. Really fast. · row: Time to Final, streaming · quote_snippet: And on its accuracy-versus-latency evaluation, we sit on the Pareto frontier
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: 达到最终转写所需时间为 0.130 秒(130ms),位于帕累托前沿。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: MAI-Voice-2.1: Seamlessly support multilingual experiences · row: 4000-listener Turing Test · quote_snippet: In a 4,000-listener Turing Test, 50.3% rated MAI-Voice as equally or more human-like than human recordings
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: 4000 名听众盲测图灵测试,50.3% 认为语音与真人录音同样或更像真人;统计综合了 MAI-Voice-2.1 与 MAI-Voice-2.1-Flash。