Inflect v2 releases two tiny local TTS models with 3.96M and 9.36M parameters
Anyphone8 · reddit · 2026-07-25
A solo developer released Inflect v2, two tiny local TTS models that generate speech from text with the waveform decoder included.
- Inflect-Nano-v2: 3.96M parameters, 15.97 MB FP32
- Inflect-Micro-v2: 9.36M parameters, 37.53 MB FP32
- Both output 24 kHz speech, run locally on CPU or CUDA, and do not require an external vocoder or hosted API.
The author says v2 is a major rebuild focused on fixing v1’s unstable timing, metallic output, weak prosody, and poor generalization. Reported results include:
- Micro: 4.395 UTMOS22, 3.99% semantic WER, 6.28× real-time CPU inference
- Nano: 4.386 UTMOS22, 4.21% semantic WER, 10.72× real-time CPU inference
- In a blind community comparison, the two models finished second and third among the tested compact voices.
Limitations remain: English-only, one fixed male voice, no voice cloning, and trouble with unusual names, abbreviations, numbers, and homographs.
The project ships with a Hugging Face playground, model cards, and GitHub code; the author says a v3 may expand voices, languages, and fine-tuning.
More from Multimodal
- FlickArtHQ adds a free 3D Stage for AI filmmaking inside Canvas — teortaxesTex · 2026-07-25
- Claude Opus 5 builds a 13-room playable penthouse on the first prompt — Silver-Chipmunk7744 · 2026-07-25
- Midjourney V8.2 becomes the default model with side-by-side prompt comparisons — LudovicCreator · 2026-07-25
- A creator is trying to make a 100-page comic with ChatGPT on a $99 Android phone — Jussi_Mulari · 2026-07-25
- A ComfyUI Gemma 4 node turns screenshots into boxed prompts for Ideogram 4 — Physical-Mission-867 · 2026-07-25
- Claude is asked for astrology brands and returns a polished shopping shortlist — ajaymehta · 2026-07-25