跳到主内容
@wquguru
精选88Hacker News Best(web_list)模型发布/更新

CactusCompute发布16.9MB端侧语音识别模型Whistle

Whistle: Speech to Text in 16.9 MB

原文
发到 X
推荐理由

端侧部署是AI落地关键场景,Whistle以极小体积和超低延迟展示了边缘推理潜力,适合关注轻量化模型与本地部署的工程团队参考。

Today we release Whistle, a speech recognition model for mobiles, wearables, robots, smart home, automotive and microcontrollers. It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle, from the same container and the same quantisation.

今天我们发布 Whistle,一款面向手机、可穿戴设备、机器人、智能家居、汽车和微控制器的语音识别模型。它由一个 16.9 MB 的文件组成,在 CPU 上运行且无依赖项,并与 Needle 使用相同的 C++ 引擎加载,来自同一个容器和相同的量化方案。

Whistle sandbox

Whistle 沙箱

Speak, and Whistle transcribes it on your device

说话,Whistle 会在你的设备上将其转录

16.9 MB · runs in this tab

16.9 MB · 在此标签页中运行

Language

语言

Keywords

关键词

Press the mic and say something.

按下麦克风并说点什么。

Up to 30 seconds in English, German, French, Spanish, Italian, Dutch or Polish. The first press downloads the 16.9 MB model, and audio never leaves your device.

支持长达 30 秒的英语、德语、法语、西班牙语、意大利语、荷兰语或波兰语。第一次按下会下载 16.9 MB 的模型,且音频永远不会离开你的设备。

Whistle does three jobs, all of them on the device:

Whistle 执行三项任务,全部在设备上完成:

  • Transcription. 16 kHz mono audio, up to 30 seconds in one pass, in English, German, French, Spanish, Italian, Dutch and Polish. The language is detected unless you name it.
  • Word timestamps. Every word with its start, end and probability, aligned from the decoder's attention.
  • Speech embedding. The encoder output, one row per 80 ms frame, without decoding a transcript.
  • 转录。16 kHz 单声道音频,单次处理最长 30 秒,支持英语、德语、法语、西班牙语、意大利语、荷兰语和波兰语。除非指定语言,否则会自动检测语言。
  • 词级时间戳。每个词及其开始时间、结束时间和概率,根据解码器的注意力机制对齐。
  • 语音嵌入。编码器输出,每 80 毫秒帧一行,无需解码出文本即可获取。

The model

模型结构

ENCODERWaveform16 kHz mono, up to 30 s = 480,000 samplesLog-mel80 bins, 25 ms window, 10 ms hop, 250-3500 Hz, normalised per channel → 3,000 framesConvolutional stem128 channels, kernel 9, three halvings → 375 frames, one per 80 msSimple Attention blocks × 8shared with Needleself-attention over every frame at once, 4 mHC lanes, Monarch Hadamard MLP in place of the FFNDECODERCross memoryK and V projected once per clip, 375 frames × 8 layers, shared by every beamLaddered Simple Attention blocks × 8shared with NeedleGQA 8q : 2kv, 48 qk / 64 v, 3-tap causal conv, engram at layers 3 and 7, width 512Gated cross attentionevery decoder layer reads the encoder: x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) VBeam search × 5length-normalised log prob, keyword bias by Aho-Corasick automatonTranscript8,192 text pieces + 7 language tokens, up to 320 of them, word times from the decoder's own attention

编码器:波形输入为 16 kHz 单声道,最长 30 秒 = 480,000 个样本;对数梅尔频谱:80 个频带,25 ms 窗口,10 ms 步长,频率范围 250-3500 Hz,按通道归一化 → 3,000 帧;卷积主干:128 个通道,核大小 9,三次减半 → 375 帧,每 80 毫秒一帧;简单注意力块 × 8:与 Needle 共享,一次性对所有帧进行自注意力计算,4 mH 通道,用 Monarch Hadamard MLP 替代前馈网络(FFN);解码器:交叉记忆:K 和 V 在每个片段中仅投影一次,375 帧 × 8 层,所有束搜索共享;分层简单注意力块 × 8:与 Needle 共享,GQA 8q:2kv,48 qk / 64 v,3-tap 因果卷积,在第 3 和第 7 层设置 engram,宽度 512;门控交叉注意力:每个解码器层读取编码器:x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) V;束搜索 × 5:长度归一化的对数概率,通过 Aho-Corasick 自动机实现关键词偏置;转录本:8,192 个文本单元 + 7 个语言标记,最多 320 个,词级时间来自解码器自身的注意力机制

Blocks marked shared run Needle's code, not a copy of it. --audio-depth selects decoder layers; the encoder always runs all eight.

标记为 shared 的部分运行的是 Needle 的代码,而非其副本。--audio-depth 参数用于选择解码器层数;编码器始终运行全部八层。

The front end. 16 kHz mono audio is framed at a 25 ms window and a 10 ms hop into 80 log-mel bins, band-limited to 250-3500 Hz and normalised per channel. Thirty seconds is 3,000 frames. A convolutional stem of 128 channels and kernel 9 halves that count three times, leaving 375 frames at one per 80 ms. Every stage after this runs at that rate, and embed returns one row per frame.

前端处理。16 kHz 单声道音频以 25 ms 窗口和 10 ms 步长分帧,形成 80 个对数梅尔频带,带宽限制在 250-3500 Hz 并按通道归一化。三十秒对应 3,000 帧。一个包含 128 个通道、核大小为 9 的卷积主干将帧数三次减半,最终剩下 375 帧,即每 80 毫秒一帧。此后所有阶段均以此速率运行,embed 输出每帧一行。

The encoder. Eight Simple Attention blocks: four mHC residual lanes and a Monarch Hadamard MLP in place of the feed-forward network, the same blocks Needle uses. The attention is not causal. A frame at 3 s attends to a frame at 12 s.

编码器。八个 Simple Attention 模块:四个 mHC 残差通道和一个 Monarch Hadamard MLP,取代前馈网络,与 Needle 使用的模块相同。注意力机制不是因果的。3秒处的帧可以关注12秒处的帧。

The decoder. Eight Laddered Simple Attention blocks at width 512, 8 query heads to 2 KV heads, 48-dimensional queries and keys, 64-dimensional values, a 3-tap causal convolution on Q, K and V, and engram lookups at layers 3 and 7 over 18,432 slots. That is Needle's block list with a different layer count.

解码器。宽度为512的八个 Laddered Simple Attention 模块,8个查询头对应2个KV头,48维查询和键,64维值,对Q、K和V进行3抽头因果卷积,并在第3层和第7层对18,432个槽位进行engram查找。这是Needle的块列表,但层数不同。

The speech-specific part is one addition per layer. Each decoder layer reads the encoder through a gated cross attention, x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) V, with a gate learned per layer and K and V taken from the clip. Those projections run once when the clip arrives, 375 frames across 8 layers, and are then held for the whole decode. Five beams therefore cost five short transcript caches, not five passes over the audio.

语音特定部分每层增加一个操作。每个解码器层通过门控交叉注意力读取编码器,x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) V,其中门控按层学习,K和V取自音频片段。这些投影在片段到达时运行一次,跨越8层的375帧,然后在整个解码过程中保持固定。因此,五个波束只需要五个简短的转录缓存,而不是五次遍历音频。

Decoding. Five beams scored by length-normalised log probability. Keyword biasing walks an Aho-Corasick automaton over the phrases you pass in, alongside the beams, and lifts their log probability as the automaton advances. The transcript is capped at 320 tokens. The vocabulary is 8,192 text pieces plus seven language tokens, one per language, so the detected language is emitted as a token rather than returned out of band.

解码。使用长度归一化的对数概率对五个波束进行评分。关键词偏置沿着你传入的短语构建Aho-Corasick自动机,与波束并行工作,并随着自动机的推进提升它们的对数概率。转录文本限制在320个token以内。词汇表包含8,192个文本单元加上七个语言token,每种语言一个,因此检测到的语言作为token输出,而不是带外返回。

The ladder is on the decoder. Every depth from 2 layers up was trained as a model of its own, and --audio-depth selects one at load time. The encoder is never sliced: all eight blocks run at every depth.

阶梯结构位于解码器上。从2层及以上的所有深度都作为独立模型进行了训练,--audio-depth参数在加载时选择其中一个。编码器从不切片:所有八个模块在每个深度都会运行。

Silence. The engine measures the clip's loudness range before the decoder starts. Below the threshold it returns an empty transcript and an empty language, and never enters the beam search.

静音。引擎在解码器启动前测量片段的响度范围。低于阈值时,它返回空转录和空语言,并且从不进入波束搜索。

Benchmarks

基准测试

WhistleWhisper baseMoonshine tiny v2

Word error rate, lower is better. A missing bar is a benchmark that model's authors never published: Moonshine is English only, and Whisper reports no SPGISpeech, Earnings-22 or AMI cleaned. Whisper's AMI figure is AMI-IHM, a different subset from the AMI the other two report.

词错误率,越低越好。缺失的条形图表示该模型的作者从未发布过相关基准:Moonshine仅支持英语,而Whisper未报告SPGISpeech、Earnings-22或AMI cleaned的数据。Whisper的AMI数据是AMI-IHM,与其他两个报告的AMI子集不同。

Whistle is ahead on LibriSpeech test-clean and test-other, on SPGISpeech, on Earnings-22 and on the FLEURS average. Whisper base is ahead on TED-LIUM, on AMI and on the MLS average, at 145.3 MB against 16.9.

Whistle在LibriSpeech test-clean和test-other、SPGISpeech、Earnings-22以及FLEURS平均值上领先。Whisper base在TED-LIUM、AMI以及MLS平均值上领先,大小为145.3 MB,而Whistle为16.9 MB。

Size

大小

megabytes, smaller is better

兆字节,越小越好

Whistle16.9 MB

Whistle 16.9 MB

Whisper base145.3 MB

Whisper base 145.3 MB

Moonshine tiny v241.9 MB

Moonshine tiny v2 41.9 MB

Time to first token

首个token的时间

milliseconds, smaller is better

毫秒,越小越好

Whistle11.1 ms

Whistle 11.1 ms

Whisper base73.2 ms

Whisper base 73.2 ms

Moonshine tiny v222.8 ms

Moonshine tiny v222.8 毫秒

Decode

解码

tokens per second, larger is better

每秒 token 数,越大越好

Whistle1,319/s

Whistle 1,319/s

Whisper base266/s

Whisper base 266/s

Moonshine tiny v2262/s

Moonshine tiny v2 262/s

Ten seconds of audio on an Apple M4 Pro CPU. Bars are scaled within each panel.

在 Apple M4 Pro CPU 上处理十秒音频。各面板内的条形图按比例缩放。

Each model ran on its official runtime at its defaults: Whistle's C++ engine at 5 beams, openai-whisper, and moonshine-voice non-streaming over whole audio. Time to first token is audio in to first token. Decode is tokens divided by the wall time after it, so the encoder is not counted twice. Whisper pads every input to 30 seconds, so its time to first token is flat across clip lengths. Whistle's tracks the clip: 5.9 ms at 5 seconds, 11.1 ms at 10, 36.3 ms at 30.

每个模型均在其官方运行时以默认配置运行:Whistle 的 C++ 引擎使用 5 个束(beams),openai-whisper,以及 moonshine-voice 对整个音频进行非流式处理。首 token 时间是从音频输入到生成第一个 token 的时间。解码速度是 token 数除以随后的墙钟时间,因此编码器不会被重复计算。Whisper 会将每个输入填充至 30 秒,因此其首 token 时间在不同片段长度下保持恒定。Whistle 则跟踪片段时长:5 秒时为 5.9 毫秒,10 秒时为 11.1 毫秒,30 秒时为 36.3 毫秒。

Word error rates are scored with the Whisper normalizers. Whistle's are measured over 86,174 utterances. Whisper's and Moonshine's are the figures their authors published, from the multilingual checkpoints rather than the English-only ones. No test audio appears in Whistle's training or validation data, verified by comparing audio checksums and speaker IDs across every reported test set.

词错误率使用 Whisper 归一化器进行评分。Whistle 的数据基于 86,174 条话语测量。Whisper 和 Moonshine 的数据为其作者发布的指标,来自多语言检查点而非仅英语检查点。经比对所有报告测试集中的音频校验和与说话人 ID,确认 Whistle 的训练或验证数据中未包含任何测试音频。

One engine, three ways to load it

一个引擎,三种加载方式

needle_load reads whichever model a .cact file holds, so the same binary does speech, text, or both:

needle_load 读取 .cact 文件中的任一模型,因此同一二进制文件可处理语音、文本或两者兼有:

代码 · 3 行
needle --model whistle.cact --audio clip.wav
needle --model needle3.cact --tools tools.json --prompt "turn off the kitchen lights"
needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav

On the third line needle_complete takes the clip directly. The engine transcribes it, answers the transcript against your tools, and returns one JSON object with the calls and the speech fields, the speech ones prefixed audio_. No transcript is handled by the caller.

第三行 needle_complete 直接接收音频片段。引擎对其进行转录,根据你的工具对转录结果进行问答,并返回一个包含 calls 和 speech 字段的 JSON 对象,其中 speech 字段以 audio_ 为前缀。转录内容由调用方处理。

代码 · 4 行
{"function_calls":[{"name":"set_lights","arguments":{"room":"kitchen","on":false}}],
 "confidence":0.94,
 "audio_text":"turn off the kitchen lights",
 "audio_language":"en"}

Get started

开始使用

代码 · 1 行
pip install cactus-needle
代码 · 3 行
import needle
print(needle.transcribe("clip.wav")["text"])
# turn off the kitchen lights

A 16 kHz WAV or raw samples need nothing beyond the base install. Other sample rates and microphone capture need the [mic] extra, which adds soxr and sounddevice.

16 kHz WAV 文件或原始采样无需额外安装,基础安装即可。其他采样率和麦克风采集需要 [mic] 扩展包,它会添加 soxr 和 sounddevice。

Every call returns the text, the language, the milliseconds to the first token and the decoder's tokens per second after it. word_timestamps=True adds each word with its times and probability. keywords=["Siobhan", "Krzysztof"] raises the log probability of those phrases during the search. language="de" forces the language instead of detecting it. needle.Whistle() is the same model as an object, for embed(audio) or to hold one tuned .cact.

每次调用均返回文本、语言、首个 token 的毫秒数以及随后解码器的每秒 token 数。word_timestamps=True 会添加每个单词及其时间和概率。keywords=["Siobhan", "Krzysztof"] 会在搜索过程中提高这些短语的对数概率。language="de" 强制指定语言而非自动检测。needle.Whistle() 以对象形式提供相同模型,可用于 embed(audio) 或保存一个微调过的 .cact 文件。

needle whistle playground transcribes from the microphone in the terminal, and needle whistle compare runs the same clip through Whistle, Whisper and Moonshine side by side with their timings.

needle whistle playground 可在终端中从麦克风转录,needle whistle compare 则将同一段音频同时通过 Whistle、Whisper 和 Moonshine 处理,并对比它们的耗时。

Deploy

部署

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件