VoiceChat-TTS: Low-Latency Streamable TTS Supporting Mid-Utterance Barge-In
rdesh26 · x · 2026-08-17
New arXiv paper VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents (Edresson Casanova et al., 20 authors).
- Problem: most speech language models remain turn-based without real-time adaptability such as user barge-in; duplex speech models cut latency but compromise quality since ASR, interruption handling, and high-fidelity synthesis must be jointly optimized.
- Approach: driven directly by LLM text-token streams; supports explicit interruption via control tokens; produces silence with no textual input; handles mid-utterance interruptions without resetting the KV cache, while preserving modularity and speech quality.
- Areas: eess.AS / cs.CL.
More from Multimodal
- MiniMax H3 used to generate absurd anime commercials — Beginning_Tip300 · 2026-08-17
- AI Generates Avatar Based on Account, Surprisingly Includes Cherry Blossoms — SydSteyerhart · 2026-08-17
- Creator Builds San Francisco Procedurally with Grok 4.6, Phase 5 Update — Daniel_Farinax · 2026-08-17
- Hailuo H3 Generates Fire Vortex Text Effect 'IGNITE' — Mr_AllenT · 2026-08-17
- Creating paper cut-out animations with H3: full motion graphics prompt workflow revealed — techhalla · 2026-08-17
- MiniMax-H3 workflow turns one design into a playable game sprite atlas for ~$5 — victormustar · 2026-08-17