New Open-Source TTS: Scylla's Band
clockentyne · reddit · 2026-07-20
A new TTS model and inference framework named Scylla's Band has been open-sourced, complete with an Android sample app.
Key highlights from the author:
- Supports 10 voices, 7 emotion/affect parameters, and 3 languages (English, Spanish, Italian).
- On a Samsung Fold 7, latency for the first audio chunk is about 500–800ms, enabling continuous streaming output thereafter.
- The code and inference stack cover Linux, macOS, and Android, with testing on M-series Macs and Android devices.
- Instead of using eSpeak, they built their own dataset, trained a G2P model, and implemented custom C++ text normalization alongside ONNX/LiteRT inference optimizations.
The author mentions future plans to expand language support, enhance emotion control, and add an iOS sample, though no promises are made yet.
More from Embodied
- Amazon and Google sold 600M+ smart speakers, so why no AGI-era successor? — julianlehr · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- Polish developers build iPhone app that detects nearby Meta smart glasses — Low-Honeydew6483 · 2026-09-11
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11