Tacit-TTS: transcript-free voice cloning 10x faster than IndexTTS2

Jian Chen · hf · 2026-10-02

Tacit-TTS is an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. It replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, adds training-free acoustic length estimation, and accelerates flow-matching via ReFlow distillation. On two English and two Mandarin datasets it matches IndexTTS2's zero-shot quality while generating speech over 10x faster for utterances longer than 5 seconds. Transcript-free conditioning also enables cross-lingual and non-lexical references, validated on eight other languages, infant babble, and synthetic gibberish where transcript-dependent systems degrade due to unreliable ASR.

Original post →

More from Multimodal

Multimodal channel →