Soniox TTS v2发布:60+语言,0.7美元/小时
Soniox TTS v2 is one of those models you need to actually hear.
Soniox TTS v2 is one of those models you need to actually hear.
Soniox TTS v2 是那种你需要亲耳听听的模型之一。
They just launched TTS v2, a text-to-speech model that speaks 60+ languages from one mode.
他们刚刚发布了 TTS v2,一个文本转语音模型,支持 60 多种语言,只需一个模型。
Really premium voice quality at a dramatically lower price ($ 0.70-per-generated-hour).
真正高级的语音质量,价格却大幅降低(每生成小时 0.70 美元)。
> One model replaces most of the voice-stack plumbing, since expression control, cloning, language mixing, pronunciation precision and streaming sit in the same system instead of 3 stitched-together vendors.
> 一个模型取代了大部分语音栈的管道,因为表达控制、克隆、语言混合、发音精确性和流式传输都在同一个系统中,而不是由三个拼凑的供应商提供。
> Voice performance becomes programmable. Because, audio tags allow developers to direct emotion, delivery, and vocal reactions throughout the text instead of selecting one fixed style for the entire passage. like [whispering], [excited] or [laughing] straight into the text,
> 语音性能变得可编程。因为音频标签允许开发者在文本中直接指定情感、表达方式和声音反应,而不是为整个段落选择一种固定风格。比如在文本中直接加入 [耳语]、[兴奋] 或 [大笑]。
> Soniox TTS v2 is built for live agents, and voice agents require more than natural-sounding audio.
> Soniox TTS v2 是为实时代理构建的,而语音代理需要的不仅仅是听起来自然的音频。
Speech must begin quickly, remain synchronized with the conversation, stop immediately when the user interrupts, and avoid repeating text that has already been spoken.
语音必须快速开始,与对话保持同步,在用户打断时立即停止,并避免重复已经说过的文本。
Interrupt a normal voice agent and it has no clue how much you actually heard, so it repeats itself or skips ahead.
打断一个普通的语音代理,它不知道你实际听到了多少,所以它会重复自己或跳过。
TTS v2 timestamps every character down to the millisecond, so it knows the exact word you cut it off at and carries on from there.
TTS v2 为每个字符打上毫秒级的时间戳,所以它知道你被切断的确切单词,并从中继续。
> Cloning from seconds of audio, keeping accent, rhythm and personality, with noise removed from the source clip first, so a phone recording still works. That voice holds its identity across all 60+ languages and switches language mid-sentence.
> 从几秒钟的音频中克隆,保留口音、节奏和个性,并首先从源剪辑中去除噪音,因此电话录音仍然有效。该声音在 60 多种语言中保持其身份,并在句子中途切换语言。
> Precision where realistic models usually break, on codes like 7Q4M9B, prices, email addresses and medical terms.
> 在现实模型通常失败的地方保持精确,如代码 7Q4M9B、价格、电子邮件地址和医学术语。
> Architecture, audio codec and inference engine are all in-house, which is what makes the price possible. Live as tts-rt-v2 in the US, Europe and Japan, backward compatible with v1.
> 架构、音频编解码器和推理引擎都是自研的,这使得价格成为可能。以 tts-rt-v2 形式在美国、欧洲和日本上线,向后兼容 v1。
🧵 1.
🧵 1.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力