Tiny 3.96M-parameter TTS now comes with voice and language fine-tuning
b111ue · reddit · 2026-07-27
The creator of Inflect v2 says you can now fine-tune the tiny TTS model on your own voice or language.
Model sizes
- Inflect-Nano-v2: 3,966,721 parameters, 15.97 MB FP32
- Inflect-Micro-v2: 9,356,513 parameters, 37.53 MB FP32
Micro is the better-sounding variant, and the author links the Hugging Face release for it.
What the new toolkit does
- accepts single-speaker recordings and transcripts
- warm-starts Nano or Micro
- resumes training
- inspects held-out output
- exports to PyTorch or ONNX
The author says they tested the full path on Nano, including a real CUDA training step, save/resume, strict PyTorch loading, and ONNX Runtime parity.
Open question
The pipeline works, but cross-language transfer is still unproven. The author has not yet produced a non-English release-quality model and says each run currently creates one fixed voice for one configured language. They’re asking others what usually breaks first when adapting small TTS models to a new language.
Related event: Inflect-Nano-v2: A 3.96M Parameter Fine-Tunable TTS Model(2 posts)→
More from Multimodal
- Screenshot any character, one prompt recreates it — iamaliveix · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- ComfyUI trick: aux preprocessor + Qwen transfers poses across characters with one prompt — Acceptable-Work8202 · 2026-09-23
- Same portrait prompt across Midjourney V6.1, V7 and V8.2: do older models look better? — tisch_eins · 2026-09-23
- Testing AI character consistency across a 20-image travel sequence — SiennaVaire · 2026-09-23
- Midjourney v8.2 Faces: New Portrait Generation Samples Shared — azed_ai · 2026-09-23