Cactus releases Whistle: a 16.9MB speech-to-text model that beats Whisper base on CPU with 6x speed
ycombinator · x · 2026-10-03
Cactus open-sourced Whistle, a speech recognition model in a single 16.9MB file that runs dependency-free on CPU, sharing the same C++ engine and quantization as their on-device foundation model Needle.
Key facts:
- Performance: 9x smaller than Whisper base, 6x faster, and mostly beats it
- Languages: English, German, French, Spanish, Italian, Dutch, Polish, with automatic language detection
- Three modes: transcription (16kHz mono, up to 30s per pass), word-level timestamps (start/end/probability), and speech embeddings (one row per 80ms frame, no decoding needed)
- Low latency: first token in 11ms; loads alongside Needle so one binary turns a clip straight into tool calls
- Target: mobiles, wearables, robots, smart home, automotive, microcontrollers — audio never leaves the device
A browser-based sandbox lets you try it directly.
More from Infra
- Debate: Nvidia's Moat Is Its Software Stack, Not Just Hardware Lock-In — QuintinPope5 · 2026-10-03
- Cerebras CEO Explains Why Wafer-Scale SRAM Beats GPU HBM by 2500x in LLM Inference — rohanpaul_ai · 2026-10-03
- Cerebras CEO Flexes His Own 42MW 13.8kV Power Generator for AI Compute — dunkhippo33 · 2026-10-03
- Is Nvidia's CUDA Moat Crumbling? DeepSeek Made Ascend Viable, Argue Researchers — QuintinPope5 · 2026-10-03
- Prime Intellect unveils Prime Inference stack after serving trillions of tokens for RL — xeophon · 2026-10-03
- Sam Altman confirms deep OpenAI-Cerebras partnership pushing inference speed frontiers — sama · 2026-10-03