Inflect v2 releases two tiny local TTS models with 3.96M and 9.36M parameters

Anyphone8 · reddit · 2026-07-25

A solo developer released Inflect v2, two tiny local TTS models that generate speech from text with the waveform decoder included.

The author says v2 is a major rebuild focused on fixing v1’s unstable timing, metallic output, weak prosody, and poor generalization. Reported results include:

Limitations remain: English-only, one fixed male voice, no voice cloning, and trouble with unusual names, abbreviations, numbers, and homographs.

The project ships with a Hugging Face playground, model cards, and GitHub code; the author says a v3 may expand voices, languages, and fine-tuning.

Original post →

More from Multimodal

Multimodal channel →