New Open-Source TTS: Scylla's Band
clockentyne · reddit · 2026-07-20
A new TTS model and inference framework named Scylla's Band has been open-sourced, complete with an Android sample app.
Key highlights from the author:
- Supports 10 voices, 7 emotion/affect parameters, and 3 languages (English, Spanish, Italian).
- On a Samsung Fold 7, latency for the first audio chunk is about 500–800ms, enabling continuous streaming output thereafter.
- The code and inference stack cover Linux, macOS, and Android, with testing on M-series Macs and Android devices.
- Instead of using eSpeak, they built their own dataset, trained a G2P model, and implemented custom C++ text normalization alongside ONNX/LiteRT inference optimizations.
The author mentions future plans to expand language support, enhance emotion control, and add an iOS sample, though no promises are made yet.
More from Embodied
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A set of agent skills for CAD, robotics, and hardware design — earthtojake · 2026-07-21
- DIY wooden box packs 6 Intel Arc Pro B70 cards with a FreeCAD model — nick_ziv · 2026-07-21
- Snake-like robot moves on fully passive wheels and winding motion — ___Mufasaa · 2026-07-21
- Creator buys a Reachy robot and asks what to build first — dee_hw · 2026-07-21
- MW team shows Gen1 of MW-bot, a semi-humanoid home robot built for pantry storage and ceiling rails — CyberRobooo · 2026-07-21