Tiny 3.96M-parameter TTS now comes with voice and language fine-tuning

b111ue · reddit · 2026-07-27

The creator of Inflect v2 says you can now fine-tune the tiny TTS model on your own voice or language.

Model sizes

Micro is the better-sounding variant, and the author links the Hugging Face release for it.

What the new toolkit does

The author says they tested the full path on Nano, including a real CUDA training step, save/resume, strict PyTorch loading, and ONNX Runtime parity.

Open question

The pipeline works, but cross-language transfer is still unproven. The author has not yet produced a non-English release-quality model and says each run currently creates one fixed voice for one configured language. They’re asking others what usually breaks first when adapting small TTS models to a new language.

Original post →

More from Multimodal

Multimodal channel →