微软发布MAI-Transcribe-2-Streaming实时语音转文字模型
Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis
Agent开发者的关键基础设施更新,首字延迟与准确率双优,直接利好实时语音交互场景,建议接入压测对比现有方案。
Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming speech-to-text (STT) model. It launched on October 1, 2026, alongside 2 text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks it #1 of 38 models for final and first partial transcript accuracy. The model targets voice agents, live captions and dictation, where latency decides the experience.
Microsoft AI 发布了 MAI-Transcribe-2-Streaming,这是其首款流式语音转文本(STT)模型。该模型于 2026 年 10 月 1 日与两款文本转语音模型 MAI-Voice-2.1 和 MAI-Voice-2.1-Flash 一同推出。Artificial Analysis 在最终转录和首个部分转录的准确率方面,将其列为 38 款模型中的第 1 名。该模型面向语音代理、实时字幕和听写场景,在这些场景中延迟决定了用户体验。
What Microsoft Shipped
微软发布了什么
MAI-Transcribe-2-Streaming is the real-time sibling of the batch MAI-Transcribe-2, released in September. It transcribes 60 languages with automatic, continuous language detection. Audio streams in continuously, and text streams back while the speaker is still talking.
MAI-Transcribe-2-Streaming 是 9 月发布的批量处理模型 MAI-Transcribe-2 的实时版本。它支持 60 种语言的自动连续语言检测。音频持续流入,而在说话者仍在讲话时,文本便已流式返回。
The model emits its first hypotheses, called partials, just over 100ms after receiving audio. It revises those partials as context arrives, then commits a stable final transcript. An agent can therefore start reasoning or calling tools mid-sentence. Microsoft team states its internal tests show words appearing 2x faster than its closest competitor.
该模型在接收音频后仅 100 多毫秒即发出首个假设结果,称为部分转录(partials)。随着上下文信息的到来,它会修订这些部分转录,随后提交稳定的最终转录结果。因此,代理可以在句子中途开始推理或调用工具。微软团队表示,其内部测试显示,单词出现速度比其最接近的竞争对手快 2 倍。
What Artificial Analysis Measured
Artificial Analysis 测量了什么
The AA-WER Streaming index uses about 8 hours of audio. The mix is AA-AgentTalk (50%), VoxPopuli (25%) and Earnings22 (25%). Latency is timed from the end of speech, as detected by SileroVAD.
AA-WER Streaming 索引使用了约 8 小时的音频。数据混合比例为 AA-AgentTalk(50%)、VoxPopuli(25%)和 Earnings22(25%)。延迟从语音结束时刻开始计时,由 SileroVAD 检测确定。
- Final transcript: 2.5% WER at 0.13s after end of speech, #1 of 38 models.
- First partial transcript: 2.5% WER at 0.12s after end of speech, also #1.
- Runners-up: Grok Voice Transcribe 2.0 at 2.7% and 0.49s. Muse Voice Transcribe at 3.1% and 0.16s.
- Not the fastest: Cartesia Ink-2 (external endpoints) returns finals in 0.07s, but at 4.0% WER.
- 最终转录:语音结束后 0.13 秒时的 WER 为 2.5%,在 38 款模型中排名第 1。
- 首个部分转录:语音结束后 0.12 秒时的 WER 为 2.5%,同样排名第 1。
- 亚军:Grok Voice Transcribe 2.0 的 WER 为 2.7%,延迟为 0.49 秒;Muse Voice Transcribe 的 WER 为 3.1%,延迟为 0.16 秒。
- 并非最快:Cartesia Ink-2(外部端点)在 0.07 秒内返回最终结果,但 WER 为 4.0%。
The first partial is as accurate as the final transcript. That matters for agents that act before the speaker finishes. Microsoft also places the model on the accuracy versus latency Pareto frontier.
首个部分转录的准确率与最终转录相当。这对于在说话者说完之前就采取行动的智能体至关重要。微软还将该模型置于准确率与延迟的帕累托前沿上。
Pricing
定价
MAI-Transcribe-2-Streaming costs $0.54 per hour of audio. This is an introductory price through the end of 2026. Artificial Analysis normalizes it to $9.00 per 1,000 minutes. Batch MAI-Transcribe-2 costs $0.10 per hour. On streaming, Microsoft charges more than xAI and Meta, and roughly matches Google’s estimated rate.
MAI-Transcribe-2-Streaming 的价格为每小时音频 0.54 美元。这是截至 2026 年底的入门价格。Artificial Analysis 将其标准化为每 1000 分钟 9.00 美元。批量处理的 MAI-Transcribe-2 价格为每小时 0.10 美元。在流式处理方面,微软的收费高于 xAI 和 Meta,大致与 Google 的估算费率持平。
How Developers Integrate It
开发者如何集成
Microsoft documents 2 integration paths. The Realtime API suits apps already using an OpenAI Realtime-compatible WebSocket. The Azure Speech SDK handles connection management, retries and audio streaming. Both return intermediate and final results.
微软文档列出了两种集成路径。Realtime API 适合已经使用兼容 OpenAI Realtime 的 WebSocket 的应用程序。Azure Speech SDK 负责连接管理、重试和音频流传输。两者均返回中间结果和最终结果。
The model is also available in the MAI Playground, through Vercel and Azure Voice Live. LiveKit support is listed as coming soon.
该模型也可在 MAI Playground 中使用,通过 Vercel 和 Azure Voice Live 提供。LiveKit 支持被列为即将推出。
Microsoft pairs it with MAI-Voice-2.1-Flash for full voice loops. Flash generates 45s of audio at 150ms end-to-end latency for $15 per 1M characters. MAI-Voice-2.1 covers 23 languages and 26 locales at $22 per 1M characters.
微软将其与 MAI-Voice-2.1-Flash 配对以实现完整的语音循环。Flash 以 150ms 的端到端延迟生成 45 秒音频,每 100 万字符收费 15 美元。MAI-Voice-2.1 覆盖 23 种语言和 26 个区域设置,每 100 万字符收费 22 美元。
Comparison: MAI-Transcribe-2-Streaming vs Closest Streaming Competitors
对比:MAI-Transcribe-2-Streaming 与最接近的流式竞争对手
| Feature | MAI-Transcribe-2-Streaming | Grok Voice Transcribe 2.0 | Muse Voice Transcribe | Gemini 3.5 Transcribe Live |
|---|---|---|---|---|
| Developer | Microsoft AI | xAI | Meta Superintelligence Labs | |
| Released | Oct 1, 2026 | Sep 18, 2026 | Sep 1, 2026 | Aug 26, 2026 |
| AA-WER Streaming (final) | 2.5% | 2.7% | 3.1% | 4.0% |
| Time to final | 0.13s | 0.49s | 0.16s | Not reported by AA source cited |
| Streaming price / hour | $0.54 (intro) | $0.20 | $0.18 | ~$0.54 (token-billed estimate) |
| Languages | 60, continuous auto-detect | Dozens, auto-detect, mid-recording switch | 70+ trained, 25 verified | 85+, auto-detect |
| Speaker diarization in stream | Not stated | Included in API (streaming not confirmed) | Yes, 20+ speakers | Not supported in Live mode |
| Interface | Realtime API (WebSocket) + Azure Speech SDK | WebSocket | WebSocket + file endpoint | Live API (WebSocket) |
| Open weights | No | No | No | No |
| 功能 | MAI-Transcribe-2-Streaming | Grok Voice Transcribe 2.0 | Muse Voice Transcribe | Gemini 3.5 Transcribe Live |
|---|---|---|---|---|
| 开发者 | Microsoft AI | xAI | Meta Superintelligence Labs | |
| 发布日期 | 2026 年 10 月 1 日 | 2026 年 9 月 18 日 | 2026 年 9 月 1 日 | 2026 年 8 月 26 日 |
| AA-WER 流式(最终) | 2.5% | 2.7% | 3.1% | 4.0% |
| 达到最终结果的时间 | 0.13 秒 | 0.49 秒 | 0.16 秒 | 引用 AA 来源未报告 |
| 流式价格/小时 | $0.54(入门价) | $0.20 | $0.18 | ~$0.54(按 token 计费的估算) |
| 语言 | 60 种,连续自动检测 | 数十种,自动检测,录音中途可切换 | 70+ 已训练,25 种已验证 | 85+,自动检测 |
| 流式中的说话人分离 | 未说明 | API 中包含(流式未确认) | 是,支持 20+ 位说话人 | Live 模式下不支持 |
| 接口 | Realtime API (WebSocket) + Azure Speech SDK | WebSocket | WebSocket + 文件端点 | Live API (WebSocket) |
| 开放权重 | 否 | 否 | 否 | 否 |
Sources: Artificial Analysis, 9to5Mac (Muse pricing), DataCamp (Grok pricing), The Batch (Gemini pricing). Verified October 2, 2026.
来源:Artificial Analysis、9to5Mac(Muse 定价)、DataCamp(Grok 定价)、The Batch(Gemini 定价)。于 2026 年 10 月 2 日核实。
Key Takeaways
关键要点
- #1 of 38 on AA-WER Streaming: 2.5% WER at 0.13s to final.
- First partials score the same 2.5% WER, at 0.12s.
- 60 languages with continuous automatic language detection.
- $0.54 per hour introductory price, higher than xAI and Meta.
- Public preview with no SLA and no open weights.
- AA-WER 流式排名 38 项中的第 1 位:最终 WER 为 2.5%,耗时 0.13 秒。
- 首个部分结果的 WER 同样为 2.5%,耗时 0.12 秒。
- 支持 60 种语言的连续自动语言检测。
- 每小时 0.54 美元的入门价格,高于 xAI 和 Meta。
- 公开预览版,无 SLA,且未开放权重。
Check out the Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
查看技术细节。所有功劳归于本项目的研究人员。此外,欢迎在 Twitter 上关注我们,并别忘了加入我们有 15 万+成员的 ML SubReddit 以及订阅我们的新闻通讯。等等!你在 Telegram 上吗?现在你也可以在 Telegram 上加入我们。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力