ItoTTS: two natural English voices in 4.89 MB running on a $5 ESP32-S3
Significant-Price695 · reddit · 2026-10-05
The Lokutor team released ItoTTS, a streaming TTS engine for the ESP32-S3 with two natural English voices, 24 kHz audio, and just 4.89 MB of weights per voice — aiming to give local LLMs a voice on a $5 chip.
In automatic UTMOS evaluation on eight held-out sentences, Ito scored 4.46, nearly matching its StyleTTS 2 teacher (4.49) and clearly beating ESP32-compatible alternatives like sanoTTS amy (3.98). The team cautions it's a small auto-eval, phoneme conversion still runs on the host, and physical-board speed is unmeasured. Code is GPLv3; weights are CC BY-NC-SA 4.0 (free non-commercial, commercial use by agreement). Code, weights, and a demo are all available.
More from Infra
- Hugging Face ships Datasets 5.1 with Vortex, Harbor RL and five bio formats — lhoestq · 2026-10-05
- boat launches: $20/mo full-VM sandboxes for AI agents with 2K concurrent instances — Scobleizer · 2026-10-05
- Nvidia guides to 70% revenue growth after four blowout quarters, bears stuck in the weeds — firstadopter · 2026-10-05
- Qwen3.8-Flash-Next (125B) runs at 59 tok/s on a single Strix Halo mini PC, engine open-sourced — Yaniss916 · 2026-10-05
- Red Hat AI ships NVFP4 quantized Qwen3.8-Flash-Next: MoE experts in FP4, vLLM-ready — huggingface · 2026-10-05
- Why Your RAM Is So Expensive: A Viral Thread Points to the AI Memory Squeeze — TheMoonMidas · 2026-10-05