Tacit-TTS: transcript-free voice cloning 10x faster than IndexTTS2
Jian Chen · hf · 2026-10-02
Tacit-TTS is an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. It replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, adds training-free acoustic length estimation, and accelerates flow-matching via ReFlow distillation. On two English and two Mandarin datasets it matches IndexTTS2's zero-shot quality while generating speech over 10x faster for utterances longer than 5 seconds. Transcript-free conditioning also enables cross-lingual and non-lexical references, validated on eight other languages, infant babble, and synthetic gibberish where transcript-dependent systems degrade due to unreliable ASR.
More from Multimodal
- Magnific shares six Ideogram 4.5 posters with full prompt iteration threads — charis_ai · 2026-10-02
- Invoke 7 PR opens after 3 months and 320 PRs: new canvas engine, projects, video with audio — joshwcorbett · 2026-10-02
- Full Cinematic Prompt Breakdown: A Mountain-Scale Storm Dragon in Four Shots — ArjanDoge · 2026-10-02
- BFL opens free trials of FLUX Tools, commercial weights available on request — bfl_ai · 2026-10-02
- FLUX 3 Image just released, per early reports — chrisfirst · 2026-10-02
- MiniMax H3 recreates the 1991 cartoon Doug — and it turned out better than expected — Certain_Potato_4509 · 2026-10-02