CTC-TTS paper replaces MFA with CTC alignment for low-latency dual-streaming TTS

kastnerkyle · x · 2026-09-13

A new arXiv paper proposes CTC-TTS, an LLM-based dual-streaming text-to-speech system. It replaces pipeline-heavy MFA forced alignment with a CTC-based neural aligner and introduces a bi-word interleaving strategy instead of fixed-ratio text-speech token interleaving.

Two variants are designed:

Experiments show CTC-TTS outperforms fixed-ratio interleaving and MFA-based baselines on streaming synthesis and zero-shot tasks. The paper was accepted at INTERSPEECH 2026.

Original post →

More from Multimodal

Multimodal channel →