Oído: open-source 13M-param speech model beats Whisper-tiny on a $5 ESP32-S3
Significant-Price695 · reddit · 2026-09-30
The Lokutor team open-sourced Oído, a speech recognition model based on NVIDIA Conformer-CTC Small (13M params, int8) that runs on a $5 ESP32-S3 with just 8 MB PSRAM, no GPU or NPU needed. It scores 3.7/8.2 WER on LibriSpeech vs 6.3/15.9 for Whisper tiny.en on a laptop, and 8.4 vs 12.1 mean WER under real-world noise and reverb (DEMAND). A livedemo.py lets you test the exact chip arithmetic with your laptop mic.
More from Infra
- DeepSeek Partners With Huawei on Chip Software to Challenge Nvidia — pstAsiatech · 2026-09-30
- a16z: tech drove 76% of S&P 500 earnings growth, and even A100 rental rates are climbing — a16z Newsletter · 2026-09-30
- TensorFlow Taught a Generation of Us ML: Moroney's Take on Picking the Right Tool — lmoroney · 2026-09-30
- Open-source Mac meeting notetaker runs 5 local models in 6.8 GB, notes in 33s — stevyhacker · 2026-09-30
- Self-hosting full GLM 5.3 takes an 8-GPU rig, roughly $136k at today's prices — TheZachMueller · 2026-09-30
- 1GW of compute yields ~$40B annual profit, equal to 1 million Tesla robotaxis, argues analyst — JOBhakdi · 2026-09-30