Whistle: an open 16.9 MB speech-to-text model that runs on CPU with 11 ms first token in 7 languages
solyarisoftware · x · 2026-10-09
Cactus released Whistle, an open-source speech recognition model that ships as a single 16.9 MB file and runs dependency-free on CPUs across phones, wearables, robots, smart home, automotive, and microcontrollers.
On-device capabilities:
- Transcription: 16 kHz mono audio, up to 30 seconds per pass, in English, German, French, Spanish, Italian, Dutch, and Polish, with automatic language detection
- Word timestamps: start/end times and probabilities per word, aligned from decoder attention
- Speech embeddings: one encoder row per 80 ms frame without decoding a transcript
Technically, 25 ms windows with a 10 ms hop produce 80 log-mel bins; a convolutional stem (128 channels, kernel 9) downsamples 3,000 frames to 375. Whistle shares the same C++ engine, container, and quantization as Cactus's Needle small foundation models (8–29 MB), so both load side by side to turn audio directly into tool calls. A browser sandbox lets users try it locally without audio leaving the device.
More from Models
- OpenAI posts 719 math proofs, two Western open-weight models debut, safety lead quits — Last Week in AI · 2026-10-09
- Open-Source NSFW Classifier Blue-Eye Hits 88.9%, Beats AWS and Google Vision — Mundane_Toe_8074 · 2026-10-09
- 'Just a wrapper' is a lazy criticism: Genspark post-trained MiniMax M3 for slides — polopat96 · 2026-10-09
- Every model looked bad in my eval — the bug was my answer key, not the models — jgarg27 · 2026-10-09
- Ai2: the next Olmo model is already training, fully open-model commitment unchanged — sewon__min · 2026-10-09
- Nous Research's typo 'Hemres Agnet' looks like a Hermes Agent teaser — Teknium · 2026-10-09