OpenAI 发布两款新转录模型
Introducing gpt-transcribe and gpt-live-transcribe
Today we're introducing two new transcription models, gpt-transcribe and gpt-live-transcribe. These models give you two different ways to build with transcription. gpt-transcribe takes a completed file and returns the full transcript. gpt-live-transcribe keeps a live open connection and returns text as the audio arrives. Both models support 57 languages and they're much better at the parts of transcription that tend to break accents, multilingual speech, very short answers, names, and numbers.
Developers can also give the models a list of words, proper nouns, or code terms that are especially important to get right. For this session, I provided a prompt mentioning transcription models plus a vocabulary with terms like phishing, ARR, and A1C from cybersecurity, sales, and healthcare domains that can be easy to miss. And because they're better at handling background noise, a side conversation or ambient music is less likely to show up as speech.
So whether you're recording in a loud cafe, a busy conference hall, or next to a very chatty coworker, the transcript can stay focused on what you're actually saying. You can give the model language hints to improve accuracy, but by default it can follow multilingual speech automatically. So I can switch languages right here in the same session. Y también puedo cambiar al español en la misma sesión. El modelo sigue transcribiendo en tiempo real y ahora vuelvo al inglés.
And now I'm back to English. Same session and the transcript followed the switch in both directions. Now let's take a look at gpt-transcribe. I've uploaded a meeting recording that's about half an hour long. The model will take a little less than a minute to process it, and then the transcript is ready for whatever comes next. It's great for call archives, podcasts, and larger batch jobs where you can wait for a complete file and optimize for accuracy and throughput.
In contrast to the streaming model, which excels at captions, dictation, and voice interfaces, where latency meaningfully impacts the experience. Happy building.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力