Oído: open-source 13M-param speech model beats Whisper-tiny on a $5 ESP32-S3

Significant-Price695 · reddit · 2026-09-30

The Lokutor team open-sourced Oído, a speech recognition model based on NVIDIA Conformer-CTC Small (13M params, int8) that runs on a $5 ESP32-S3 with just 8 MB PSRAM, no GPU or NPU needed. It scores 3.7/8.2 WER on LibriSpeech vs 6.3/15.9 for Whisper tiny.en on a laptop, and 8.4 vs 12.1 mean WER under real-world noise and reverb (DEMAND). A livedemo.py lets you test the exact chip arithmetic with your laptop mic.

Original post →

More from Infra

Infra channel →