CTC-TTS paper replaces MFA with CTC alignment for low-latency dual-streaming TTS
kastnerkyle · x · 2026-09-13
A new arXiv paper proposes CTC-TTS, an LLM-based dual-streaming text-to-speech system. It replaces pipeline-heavy MFA forced alignment with a CTC-based neural aligner and introduces a bi-word interleaving strategy instead of fixed-ratio text-speech token interleaving.
Two variants are designed:
- CTC-TTS-L: token concatenation along sequence length, for higher quality;
- CTC-TTS-F: embedding stacking along feature dimension, for lower latency.
Experiments show CTC-TTS outperforms fixed-ratio interleaving and MFA-based baselines on streaming synthesis and zero-shot tasks. The paper was accepted at INTERSPEECH 2026.
More from Multimodal
- MiniMax H3 Turns a Drop of Coffee Into an Entire City in Hailuo Design Hub — umesh_ai · 2026-09-14
- "AI Can Never Produce This Quality" — Higgsfield Claps Back With "What If It Actually Can?" — _AustinCalvert_ · 2026-09-14
- Creator Chains Seedream 2.5 with Flux 3 via Luma to Make a 40-Second AI Short — LudovicCreator · 2026-09-14
- A Full Prompt Template for Generating a Photoreal Isometric ARPG Trailer with AI — techhalla · 2026-09-14
- HeyGen open-sources 20+ launch videos as plain HTML that Claude Code can edit — _AustinCalvert_ · 2026-09-14
- Wan3.0 Anime Battle Clip Impresses: Sensible Camera Cuts, Consistent Motion — andrew_n_carr · 2026-09-14