Inflect v2 releases two tiny local TTS models with 3.96M and 9.36M parameters
Anyphone8 · reddit · 2026-07-25
A solo developer released Inflect v2, two tiny local TTS models that generate speech from text with the waveform decoder included.
- Inflect-Nano-v2: 3.96M parameters, 15.97 MB FP32
- Inflect-Micro-v2: 9.36M parameters, 37.53 MB FP32
- Both output 24 kHz speech, run locally on CPU or CUDA, and do not require an external vocoder or hosted API.
The author says v2 is a major rebuild focused on fixing v1’s unstable timing, metallic output, weak prosody, and poor generalization. Reported results include:
- Micro: 4.395 UTMOS22, 3.99% semantic WER, 6.28× real-time CPU inference
- Nano: 4.386 UTMOS22, 4.21% semantic WER, 10.72× real-time CPU inference
- In a blind community comparison, the two models finished second and third among the tested compact voices.
Limitations remain: English-only, one fixed male voice, no voice cloning, and trouble with unusual names, abbreviations, numbers, and homographs.
The project ships with a Hugging Face playground, model cards, and GitHub code; the author says a v3 may expand voices, languages, and fine-tuning.
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Tencent open-sources AuK, a unified 1.5B speech generation and editing model — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11