Xiaohongshu Open-Sources dots.tts: 2B-Param Continuous AR Model for Low-Latency Voice Cloning
小红书技术REDtech · wechat · 2026-08-13
Xiaohongshu's dots team has open-sourced dots.tts, a 2B-parameter end-to-end text-to-speech model. Instead of using discrete acoustic tokens, it performs autoregressive generation directly in a continuous latent space, achieving high-fidelity voice cloning and long-term stability with significantly lower inference latency.
On the Seed-TTS-Eval benchmark, dots.tts achieves the best average content accuracy and speaker similarity among comparable systems. Using self-correction alignment and distillation, the team compressed generation to 1, 2, or 4 steps. In its dual-stream interactive mode, the model achieves a minimum time-to-first-audio of 54.4ms, making it highly suitable for real-time voice agents.
The release includes 6 checkpoints covering pre-training and low-step generation, alongside the full training, fine-tuning, distillation, and inference code under Apache 2.0. The team also introduced dots.tts.edit for precise text, emotion, and prosody editing.
Related event: Xiaohongshu Open-Sources 2B Parameter TTS Model dots.tts(2 posts)→
More from Multimodal
- Xiaomi Phone's AI Hallucinates Moon Craters on the Sun — FlorianGallwitz · 2026-08-13
- Looking for free AI 3D model generators that allow downloading models — Reasonable_Low404 · 2026-08-13
- ComfyUI img2img masks broken in latest version, user seeks help — AltruisticList6000 · 2026-08-13
- Developer releases KREA 2 film workflow with custom ComfyUI node — Lower-Cap7381 · 2026-08-13
- Grok + Blender MCP: Animating 3D Models with a Single Prompt — chongdashu · 2026-08-13
- AI Video Character Consistency: How to Keep the Same Character Across Shots? — Usedependence · 2026-08-13