Whistle: an open 16.9 MB speech-to-text model that runs on CPU with 11 ms first token in 7 languages

solyarisoftware · x · 2026-10-09

Cactus released Whistle, an open-source speech recognition model that ships as a single 16.9 MB file and runs dependency-free on CPUs across phones, wearables, robots, smart home, automotive, and microcontrollers.

On-device capabilities:

Technically, 25 ms windows with a 10 ms hop produce 80 log-mel bins; a convolutional stem (128 channels, kernel 9) downsamples 3,000 frames to 375. Whistle shares the same C++ engine, container, and quantization as Cactus's Needle small foundation models (8–29 MB), so both load side by side to turn audio directly into tool calls. A browser sandbox lets users try it locally without audio leaving the device.

Original post →

More from Models

Models channel →