Xiaohongshu Open-Sources dots.tts: 2B-Param Continuous AR Model for Low-Latency Voice Cloning

小红书技术REDtech · wechat · 2026-08-13

Xiaohongshu's dots team has open-sourced dots.tts, a 2B-parameter end-to-end text-to-speech model. Instead of using discrete acoustic tokens, it performs autoregressive generation directly in a continuous latent space, achieving high-fidelity voice cloning and long-term stability with significantly lower inference latency.

On the Seed-TTS-Eval benchmark, dots.tts achieves the best average content accuracy and speaker similarity among comparable systems. Using self-correction alignment and distillation, the team compressed generation to 1, 2, or 4 steps. In its dual-stream interactive mode, the model achieves a minimum time-to-first-audio of 54.4ms, making it highly suitable for real-time voice agents.

The release includes 6 checkpoints covering pre-training and low-step generation, alongside the full training, fine-tuning, distillation, and inference code under Apache 2.0. The team also introduced dots.tts.edit for precise text, emotion, and prosody editing.

Related event: Xiaohongshu Open-Sources 2B Parameter TTS Model dots.tts(2 posts)→

Original post →

More from Multimodal

Multimodal channel →