SpaceXAI发布Grok Voice Transcribe
SpaceXAI Releases Grok Voice Transcribe 2.0: A Speech-to-Text API Claiming 2x Accuracy Over 1.0 at $0.10 per Hour
语音转写是Agent落地的关键基础设施,这款新模型在嘈杂环境和多语言场景下的精度提升显著且定价极具竞争力,做语音交互或内容自动化的团队值得重点关注和压测。
SpaceXAI has released Grok Voice Transcribe 2.0, its newest speech-to-text (STT) model. The development team claims it to be twice as accurate as Grok Voice Transcribe 1.0 at the same price. The model targets hard audio: noisy phone lines, competing voices, local accents, and spoken credentials. It runs in batch and real-time streaming modes through the Speech to Text API.
SpaceXAI 发布了 Grok Voice Transcribe 2.0,这是其最新的语音转文本(STT)模型。开发团队声称,在相同价格下,其准确率是 Grok Voice Transcribe 1.0 的两倍。该模型针对困难音频场景:嘈杂的电话线路、竞争性的声音、地方口音以及口语化的凭证信息。它通过 Speech to Text API 以批处理和实时流模式运行。
Is it deployable? Yes, as a hosted API. It is live today under the model ID grok-voice-transcribe-2.0. SpaceXAI has not announced open weights, so self-hosting is not an option.
它可以部署吗?可以,作为托管 API。目前已在模型 ID grok-voice-transcribe-2.0 下上线。SpaceXAI 尚未宣布开放权重,因此自托管不可行。
What is Grok Voice Transcribe 2.0?
什么是 Grok Voice Transcribe 2.0?
Grok Voice Transcribe 2.0 is built on the audio foundation model behind Grok Voice. SpaceXAI team states that Grok Voice already handles tens of thousands of customer-support calls a day. It also transcribes millions of hours of video narration and runs the Grok assistant in Tesla vehicles.
Grok Voice Transcribe 2.0 构建于 Grok Voice 背后的音频基础模型之上。SpaceXAI 团队指出,Grok Voice 每天已处理数万通客户支持电话。它还转录数百万小时的视频旁白,并在特斯拉车辆中运行 Grok 助手。
The training data is live, noisy, multilingual audio recorded across diverse environments. SpaceXAI then refined the model with post-training.
训练数据是来自多样化环境录制的真实、嘈杂的多语言音频。随后,SpaceXAI 通过后训练对模型进行了优化。
Benchmarks: What SpaceXAI Reports
基准测试:SpaceXAI 的报告结果
SpaceXAI reports a first-place accuracy rank among 32 streaming models on the public Artificial Analysis leaderboard. That benchmark, AA-WER Streaming, uses about 8 hours of audio. It weights AA-AgentTalk at 50%, VoxPopuli at 25%, and Earnings22 at 25%. See the methodology for details.
SpaceXAI 报告称,在公共 Artificial Analysis 排行榜上,其在 32 个流式模型中准确率排名第一。该基准测试名为 AA-WER Streaming,使用约 8 小时的音频。其中 AA-AgentTalk 占 50%,VoxPopuli 占 25%,Earnings22 占 25%。详见方法论说明。
SpaceXAI also measures word error rate (WER) on 4 internal sets drawn from production traffic:
SpaceXAI 还测量了来自生产流量的 4 个内部数据集的词错误率(WER):
- Telephony (8 kHz): English customer support calls
- Conversational: English conversations with Grok
- Credentials: phone numbers, emails, and addresses in English
- Short phrases: voice-assistant utterances in 19 languages
- 电话通话(8 kHz):英语客户支持电话
- 对话:与 Grok 的英语对话
- 凭证信息:英语中的电话号码、电子邮件和地址
- 短短语:19 种语言的语音助手 utterances(语音片段/指令)
Version 2.0 improves on 1.0 across all 4 sets. On telephony, SpaceXAI says it leads every model the company tested. These internal results are vendor-reported and not independently reproduced.
2.0 版本在所有 4 个数据集上均优于 1.0 版本。在电话通话场景中,SpaceXAI 表示其领先于该公司测试的所有其他模型。这些内部结果为厂商报告,未经独立复现验证。
Multilingual Transcription and Language Switching
多语言转录与语言切换
The model transcribes dozens of languages and detects the language automatically. It also follows mid-recording language switches in a single pass. SpaceXAI team calls multilingual accuracy the largest improvement over 1.0.
该模型可转录数十种语言并自动检测语言。它还能够在单次处理过程中跟随录音中途的语言切换。SpaceXAI 团队称,多语言准确率是相比 1.0 版本最大的改进。
Short phrases, such as in-car commands, give a model little context to identify the language. On that set, WER drops from 20.6% to 6.8%. That works out to roughly 67% fewer word errors.
对于短短语(如车内指令),模型缺乏足够的上下文来识别语言。在该数据集上,词错误率从 20.6% 降至 6.8%,相当于词错误减少了约 67%。
The docs list 25 languages for written-form formatting of numbers, currencies, and units.
文档列出了 25 种语言,用于数字、货币和单位的书面格式化处理。
Key Features for Developers
面向开发者的关键功能
Every feature below ships in the same API:
以下所有功能均在同一 API 中提供:
- Batch and streaming: transcribe files and URLs, or stream audio over WebSocket at wss://api.x.ai/v1/stt
- Word-level timestamps: start and end times plus confidence scores for each word
- Speaker diarization: speaker labels at no additional cost
- Multichannel transcription: up to 8 channels transcribed independently
- Key term biasing: up to 100 domain terms per request, each up to 50 characters
- Text formatting: numbers, dates, currencies, phone numbers, and emails returned in written form
- Filler word removal: “um” and “uh” are removed by default
- Smart turn detection: an ML model predicts end of turn for voice agents
- 批处理与流式传输:转录文件和 URL,或通过 WebSocket 在 wss://api.x.ai/v1/stt 上流式传输音频
- 词级时间戳:每个词的起始和结束时间以及置信度分数
- 说话人分离:无需额外费用即可获取说话人标签
- 多通道转录:最多可独立转录 8 个通道
- 关键词偏置:每次请求最多支持 100 个领域术语,每个术语最多 50 个字符
- 文本格式化:数字、日期、货币、电话号码和电子邮件以书面形式返回
- 填充词移除:“um”和“uh”默认被移除
- 智能轮次检测:机器学习模型预测语音代理的轮次结束
The batch endpoint accepts files up to 500 MB across 12 audio formats. Streaming also accepts Opus at roughly 4 KB/s, versus 48 KB/s for raw PCM at 24 kHz.
批处理端点接受最大 500 MB 的文件,涵盖 12 种音频格式。流式传输也接受 Opus 格式(约 4 KB/s),而 24 kHz 原始 PCM 为 48 KB/s。
Pricing
定价
Pricing is identical to version 1.0. Batch transcription costs $0.10 per hour of audio. Streaming costs $0.20 per hour. Diarization, timestamps, and key terms are included. That equals about $1.67 and $3.33 per 1,000 minutes.
定价与 1.0 版本相同。批处理转录每小时音频成本为 0.10 美元。流式传输每小时成本为 0.20 美元。说话人分离、时间戳和关键词均包含在内。这相当于每 1,000 分钟约 1.67 美元和 3.33 美元。
How to Call It
如何调用
curl -X POST https://api.x.ai/v1/stt \
-H "Authorization: Bearer $XAI_API_KEY" \
-F model=grok-voice-transcribe-2.0 \
-F format=true \
-F language=en \
-F [email protected]curl -X POST https://api.x.ai/v1/stt \
-H "Authorization: Bearer $XAI_API_KEY" \
-F model=grok-voice-transcribe-2.0 \
-F format=true \
-F language=en \
-F [email protected]Atlassian Loom Adopts Grok Voice Transcribe 2.0
Atlassian Loom 采用 Grok Voice Transcribe 2.0
Atlassian Loom now uses Grok Voice Transcribe 2.0 to transcribe every video. Atlassian found it more accurate than its existing solution.
Atlassian Loom 现在使用 Grok Voice Transcribe 2.0 转录所有视频。Atlassian 发现其准确性高于现有解决方案。
The workflow SpaceXAI describes is record, transcribe, then code. A user records an action plan in Loom. The transcript is piped into Cursor, which makes the code updates. Sanchan Saxena, SVP of Teamwork Collection at Atlassian, described it as ‘closing the loop from context to code.’
SpaceXAI 描述的工作流程是录制、转录,然后编码。用户在 Loom 中录制行动计划。转录内容被输入到 Cursor 中,从而进行代码更新。Atlassian 团队合作副总裁 Sanchan Saxena 将其描述为‘从上下文到代码的闭环’。
Key Takeaways
关键要点
- Grok Voice Transcribe 2.0 claims 2x the accuracy of 1.0 at unchanged pricing.
- Batch costs $0.10 per audio hour and streaming costs $0.20.
- SpaceXAI reports rank 1 among 32 streaming models on Artificial Analysis.
- Short-phrase WER across 19 languages fell from 20.6% to 6.8%.
- It is API-only, so set model=grok-voice-transcribe-2.0 explicitly for now.
- Grok Voice Transcribe 2.0 声称准确率是 1.0 的两倍,且价格不变。
- 批处理每小时音频成本为 0.10 美元,流式传输为 0.20 美元。
- SpaceXAI 在 Artificial Analysis 的 32 个流式模型中排名第一。
- 19 种语言的短短语 WER 从 20.6% 降至 6.8%。
- 仅支持 API,因此目前需显式设置 model=grok-voice-transcribe-2.0。
Check out the technical details and the API docs. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
查看技术细节和 API 文档。所有功劳归于该项目的研究人员。此外,欢迎在 Twitter 上关注我们,别忘了加入我们有 15 万+成员的 ML SubReddit 并订阅我们的新闻通讯。等等!你在 Telegram 上吗?现在你也可以加入我们的 Telegram 群组。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力