Xiaohongshu Open-Sources dots.tts: 2B Continuous AR Speech Model

机器之心 · wechat · 2026-08-13

Xiaohongshu's dots team, collaborating with Shanghai Jiao Tong University, has open-sourced dots.tts, a 2B parameter end-to-end text-to-speech foundation model. Departing from the mainstream discrete token approach, dots.tts performs autoregressive generation directly in a continuous latent space, preserving richer acoustic details like tone and prosody.

On popular zero-shot voice cloning benchmarks such as Seed-TTS-Eval, dots.tts achieves SOTA in both average content accuracy and speaker similarity. The model natively supports streaming output, achieving an ultra-low audio first-package latency of 54.4ms for duplex dialogue systems. The release includes six checkpoints covering base models, self-corrected alignment, and multi-step/single-step distillation, alongside fully open training and fine-tuning code.

Original post →

More from Multimodal

Multimodal channel →