Whistle:16.9MB端侧语音识别模型发布
Whistle: speech to text in 16.9 MB
极致压缩的端侧ASR方案,16.9MB体积配合高吞吐低延迟,非常适合嵌入式与IoT场景的开发者参考落地。
Today we release Whistle, a speech recognition model for mobiles, wearables, robots, smart home, automotive and microcontrollers. It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle, from the same container and the same quantisation.
今天我们发布 Whistle,一款面向手机、可穿戴设备、机器人、智能家居、汽车和微控制器的语音识别模型。它仅包含一个 16.9 MB 的文件,在 CPU 上运行且无依赖项,并加载到与 Needle 相同的 C++ 引擎中,来自同一个容器和相同的量化方案。
Whistle sandbox
Whistle 沙箱
Speak, and Whistle transcribes it on your device
说话,Whistle 会在你的设备上将其转录为文本
16.9 MB · runs in this tab
16.9 MB · 在此标签页中运行
Language
语言
Keywords
关键词
Press the mic and say something.
按下麦克风并说点什么。
Up to 30 seconds in English, German, French, Spanish, Italian, Dutch or Polish. The first press downloads the 16.9 MB model, and audio never leaves your device.
支持长达 30 秒的英语、德语、法语、西班牙语、意大利语、荷兰语或波兰语。第一次按下会下载 16.9 MB 的模型,且音频永远不会离开你的设备。
Whistle does three jobs, all of them on the device:
Whistle 执行三项任务,全部在设备端完成:
- Transcription. 16 kHz mono audio, up to 30 seconds in one pass, in English, German, French, Spanish, Italian, Dutch and Polish. The language is detected unless you name it.
- Word timestamps. Every word with its start, end and probability, aligned from the decoder's attention.
- Speech embedding. The encoder output, one row per 80 ms frame, without decoding a transcript.
- 转录。16 kHz 单声道音频,单次处理最长 30 秒,支持英语、德语、法语、西班牙语、意大利语、荷兰语和波兰语。除非指定语言,否则会自动检测语言。
- 词级时间戳。每个词及其开始时间、结束时间和概率,由解码器的注意力机制对齐。
- 语音嵌入。编码器输出,每 80 毫秒帧对应一行,无需解码出转录文本。
The model
模型结构
ENCODERWaveform16 kHz mono, up to 30 s = 480,000 samplesLog-mel80 bins, 25 ms window, 10 ms hop, 250-3500 Hz, normalised per channel → 3,000 framesConvolutional stem128 channels, kernel 9, three halvings → 375 frames, one per 80 msSimple Attention blocks × 8shared with Needleself-attention over every frame at once, 4 mHC lanes, Monarch Hadamard MLP in place of the FFNDECODERCross memoryK and V projected once per clip, 375 frames × 8 layers, shared by every beamLaddered Simple Attention blocks × 8shared with NeedleGQA 8q : 2kv, 48 qk / 64 v, 3-tap causal conv, engram at layers 3 and 7, width 512Gated cross attentionevery decoder layer reads the encoder: x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) VBeam search × 5length-normalised log prob, keyword bias by Aho-Corasick automatonTranscript8,192 text pieces + 7 language tokens, up to 320 of them, word times from the decoder's own attention
编码器:波形输入为 16 kHz 单声道,最长 30 秒 = 480,000 个样本;对数梅尔频谱:80 个频带,25 毫秒窗口,10 毫秒步长,频率范围 250-3500 Hz,按通道归一化 → 3,000 帧;卷积主干:128 个通道,核大小 9,三次减半操作 → 375 帧,每 80 毫秒一帧;简单注意力块 × 8:与 Needle 共享,同时对每一帧进行自注意力计算,4 mH 通道,用 Monarch Hadamard MLP 替代前馈网络(FFN);解码器:交叉记忆机制,K 和 V 在每个片段中仅投影一次,375 帧 × 8 层,由所有束搜索路径共享;分层简单注意力块 × 8:与 Needle 共享,GQA 配置为 8q:2kv,48 qk / 64 v,3 抽头因果卷积,在第 3 和第 7 层引入 engram 机制,宽度为 512;门控交叉注意力:每个解码器层读取编码器输出:x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) V;束搜索 × 5:长度归一化的对数概率,通过 Aho-Corasick 自动机实现关键词偏置;转录结果:8,192 个文本单元 + 7 个语言标记,最多 320 个,词级时间来自解码器自身的注意力机制
Blocks marked shared run Needle's code, not a copy of it. --audio-depth selects decoder layers; the encoder always runs all eight.
标记为 shared 的模块运行的是 Needle 的代码,而非其副本。--audio-depth 参数用于选择解码器层数;编码器始终运行全部八层。
The front end. 16 kHz mono audio is framed at a 25 ms window and a 10 ms hop into 80 log-mel bins, band-limited to 250-3500 Hz and normalised per channel. Thirty seconds is 3,000 frames. A convolutional stem of 128 channels and kernel 9 halves that count three times, leaving 375 frames at one per 80 ms. Every stage after this runs at that rate, and embed returns one row per frame.
前端处理。16 kHz 单声道音频以 25 毫秒窗口和 10 毫秒步长进行分帧,生成 80 个对数梅尔频带,带宽限制在 250-3500 Hz 并按通道归一化。三十秒对应 3,000 帧。一个具有 128 个通道和核大小 9 的卷积主干经过三次减半操作,将帧数减少至 375 帧,即每 80 毫秒一帧。此后所有阶段均以此速率运行,嵌入输出每帧返回一行。
The encoder. Eight Simple Attention blocks: four mHC residual lanes and a Monarch Hadamard MLP in place of the feed-forward network, the same blocks Needle uses. The attention is not causal. A frame at 3 s attends to a frame at 12 s.
编码器。八个 Simple Attention 模块:四个 mHC 残差通道和一个 Monarch Hadamard MLP,取代前馈网络,与 Needle 使用的模块相同。注意力机制非因果性。3秒处的帧可以关注12秒处的帧。
The decoder. Eight Laddered Simple Attention blocks at width 512, 8 query heads to 2 KV heads, 48-dimensional queries and keys, 64-dimensional values, a 3-tap causal convolution on Q, K and V, and engram lookups at layers 3 and 7 over 18,432 slots. That is Needle's block list with a different layer count.
解码器。宽度为512的八个 Laddered Simple Attention 模块,8个查询头对应2个KV头,48维查询和键,64维值,对Q、K和V进行3抽头因果卷积,并在第3层和第7层对18,432个槽位进行engram查找。即 Needle 的块列表,但层数不同。
The speech-specific part is one addition per layer. Each decoder layer reads the encoder through a gated cross attention, x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) V, with a gate learned per layer and K and V taken from the clip. Those projections run once when the clip arrives, 375 frames across 8 layers, and are then held for the whole decode. Five beams therefore cost five short transcript caches, not five passes over the audio.
语音特定部分每层增加一个操作。每个解码器层通过门控交叉注意力读取编码器,公式为 x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) V,其中门控按层学习,K和V取自音频片段。这些投影在片段到达时运行一次,跨越8层共375帧,然后在整个解码过程中保持固定。因此,五个波束只需五个简短的转录缓存,而非五次遍历音频。
Decoding. Five beams scored by length-normalised log probability. Keyword biasing walks an Aho-Corasick automaton over the phrases you pass in, alongside the beams, and lifts their log probability as the automaton advances. The transcript is capped at 320 tokens. The vocabulary is 8,192 text pieces plus seven language tokens, one per language, so the detected language is emitted as a token rather than returned out of band.
解码。由长度归一化对数概率评分的五个波束。关键词偏置沿传入的短语构建 Aho-Corasick 自动机,与波束并行运行,并随着自动机推进提升其对数概率。转录文本限制为320个token。词表包含8,192个文本单元加上七个语言token(每种语言一个),因此检测到的语言作为token输出,而非带外返回。
The ladder is on the decoder. Every depth from 2 layers up was trained as a model of its own, and --audio-depth selects one at load time. The encoder is never sliced: all eight blocks run at every depth.
阶梯结构位于解码器。从2层及以上的所有深度均作为独立模型训练,--audio-depth 参数在加载时选择其中一个。编码器从不切片:所有八个模块在每个深度下均运行。
Silence. The engine measures the clip's loudness range before the decoder starts. Below the threshold it returns an empty transcript and an empty language, and never enters the beam search.
静音。引擎在解码器启动前测量片段的响度范围。低于阈值时,返回空转录和空语言,且不进入波束搜索。
Benchmarks
基准测试
WhistleWhisper baseMoonshine tiny v2
Word error rate, lower is better. A missing bar is a benchmark that model's authors never published: Moonshine is English only, and Whisper reports no SPGISpeech, Earnings-22 or AMI cleaned. Whisper's AMI figure is AMI-IHM, a different subset from the AMI the other two report.
词错误率,越低越好。缺失的条形图表示该模型的作者从未发布过相关基准:Moonshine 仅支持英语,且 Whisper 未报告 SPGISpeech、Earnings-22 或 AMI cleaned 的结果。Whisper 的 AMI 数据是 AMI-IHM,与其他两个模型报告的 AMI 子集不同。
Whistle is ahead on LibriSpeech test-clean and test-other, on SPGISpeech, on Earnings-22 and on the FLEURS average. Whisper base is ahead on TED-LIUM, on AMI and on the MLS average, at 145.3 MB against 16.9.
Whistle 在 LibriSpeech test-clean 和 test-other、SPGISpeech、Earnings-22 以及 FLEURS 平均值上领先。Whisper base 在 TED-LIUM、AMI 和 MLS 平均值上领先,大小为145.3 MB,而 Whistle 为16.9 MB。
Size
大小
megabytes, smaller is better
兆字节,越小越好
Whistle16.9 MB
Whistle 16.9 MB
Whisper base145.3 MB
Whisper base 145.3 MB
Moonshine tiny v241.9 MB
Moonshine tiny v2 41.9 MB
Time to first token
首token时间
milliseconds, smaller is better
毫秒,越小越好
Whistle11.1 ms
Whistle 11.1 ms
Whisper base73.2 ms
Whisper base 73.2 ms
Moonshine tiny v222.8 ms
Moonshine tiny v2 222.8 毫秒
Decode
解码
tokens per second, larger is better
每秒 token 数,越大越好
Whistle1,319/s
Whistle 1,319/s
Whisper base266/s
Whisper base 266/s
Moonshine tiny v2262/s
Moonshine tiny v2 262/s
Ten seconds of audio on an Apple M4 Pro CPU. Bars are scaled within each panel.
在 Apple M4 Pro CPU 上处理十秒音频。各面板内的柱状图按比例缩放。
Each model ran on its official runtime at its defaults: Whistle's C++ engine at 5 beams, openai-whisper, and moonshine-voice non-streaming over whole audio. Time to first token is audio in to first token. Decode is tokens divided by the wall time after it, so the encoder is not counted twice. Whisper pads every input to 30 seconds, so its time to first token is flat across clip lengths. Whistle's tracks the clip: 5.9 ms at 5 seconds, 11.1 ms at 10, 36.3 ms at 30.
每个模型均在其官方运行时以默认配置运行:Whistle 的 C++ 引擎使用 5 个束搜索,openai-whisper 和 moonshine-voice 对整段音频进行非流式处理。首 token 时间指从音频输入到生成第一个 token 的时间。解码速度为 token 数除以随后的墙钟时间,因此编码器不被重复计算。Whisper 会将所有输入填充至 30 秒,因此其首 token 时间在不同片段长度下保持恒定。Whistle 则跟踪片段长度:5 秒时为 5.9 毫秒,10 秒时为 11.1 毫秒,30 秒时为 36.3 毫秒。
Word error rates are scored with the Whisper normalizers. Whistle's are measured over 86,174 utterances. Whisper's and Moonshine's are the figures their authors published, from the multilingual checkpoints rather than the English-only ones. No test audio appears in Whistle's training or validation data, verified by comparing audio checksums and speaker IDs across every reported test set.
词错误率使用 Whisper 归一化器进行评分。Whistle 的数据基于 86,174 条话语测量得出。Whisper 和 Moonshine 的数据为其作者发布的数值,来自多语言检查点而非仅英语版本。通过比对所有报告测试集中的音频校验和与说话人 ID,已验证 Whistle 的训练或验证数据中不包含任何测试音频。
One engine, three ways to load it
一个引擎,三种加载方式
needle_load reads whichever model a .cact file holds, so the same binary does speech, text, or both:
needle_load 读取 .cact 文件中的模型,因此同一二进制文件可处理语音、文本或两者兼有:
needle --model whistle.cact --audio clip.wav
needle --model needle3.cact --tools tools.json --prompt "turn off the kitchen lights"
needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wavOn the third line needle_complete takes the clip directly. The engine transcribes it, answers the transcript against your tools, and returns one JSON object with the calls and the speech fields, the speech ones prefixed audio_. No transcript is handled by the caller.
第三行中 needle_complete 直接接收音频片段。引擎对其进行转录,根据工具回答转录内容,并返回包含 calls 和 speech 字段的单个 JSON 对象,其中语音相关字段以 audio_ 为前缀。转录内容由调用方处理。
{"function_calls":[{"name":"set_lights","arguments":{"room":"kitchen","on":false}}],
"confidence":0.94,
"audio_text":"turn off the kitchen lights",
"audio_language":"en"}Get started
入门指南
pip install cactus-needleimport needle
print(needle.transcribe("clip.wav")["text"])
# turn off the kitchen lightsA 16 kHz WAV or raw samples need nothing beyond the base install. Other sample rates and microphone capture need the [mic] extra, which adds soxr and sounddevice.
16 kHz WAV 文件或原始采样无需额外安装,仅需基础环境。其他采样率和麦克风捕获需要安装 [mic] 扩展,它将添加 soxr 和 sounddevice。
Every call returns the text, the language, the milliseconds to the first token and the decoder's tokens per second after it. word_timestamps=True adds each word with its times and probability. keywords=["Siobhan", "Krzysztof"] raises the log probability of those phrases during the search. language="de" forces the language instead of detecting it. needle.Whistle() is the same model as an object, for embed(audio) or to hold one tuned .cact.
每次调用均返回文本、语言、首个 token 的毫秒数以及随后的解码器每秒 token 数。设置 word_timestamps=True 会添加每个单词及其时间和概率。keywords=["Siobhan", "Krzysztof"] 会在搜索过程中提高这些短语的对数概率。language="de" 强制指定语言而非自动检测。needle.Whistle() 将模型作为对象返回,可用于 embed(audio) 或保存经过微调的 .cact 文件。
needle whistle playground transcribes from the microphone in the terminal, and needle whistle compare runs the same clip through Whistle, Whisper and Moonshine side by side with their timings.
needle whistle playground 可在终端中从麦克风转录音频,needle whistle compare 则将同一段音频同时送入 Whistle、Whisper 和 Moonshine 进行对比,并显示各自的处理时间。
Deploy
部署
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力