Xiaohongshu Open-Sources dots.tts: 2B Continuous AR Speech Model
机器之心 · wechat · 2026-08-13
Xiaohongshu's dots team, collaborating with Shanghai Jiao Tong University, has open-sourced dots.tts, a 2B parameter end-to-end text-to-speech foundation model. Departing from the mainstream discrete token approach, dots.tts performs autoregressive generation directly in a continuous latent space, preserving richer acoustic details like tone and prosody.
On popular zero-shot voice cloning benchmarks such as Seed-TTS-Eval, dots.tts achieves SOTA in both average content accuracy and speaker similarity. The model natively supports streaming output, achieving an ultra-low audio first-package latency of 54.4ms for duplex dialogue systems. The release includes six checkpoints covering base models, self-corrected alignment, and multi-step/single-step distillation, alongside fully open training and fine-tuning code.
More from Multimodal
- Xiaomi's New Framework Combines Physics Simulation and Diffusion for Video Dereflection — xiaomi-research · 2026-08-13
- NeuPAT: Neuron-aware Tuning Preserves Language Skills in Multimodal LLMs — CASIA-IVA-Lab · 2026-08-13
- Study Reveals Lack of Causal Effectiveness in Multimodal LLM Visual Tool-Use — Zhiheng Wang · 2026-08-13
- Testing H3 Minimax on Dual RTX 4080s: Hardware Isn't the Bottleneck Anymore, Creativity Is — writingdeveloper · 2026-08-13
- Structured JSON Prompts for Video Generation: Porting Sora Prompts Directly to Minimax — ajrss2009 · 2026-08-13
- ComfyUI v0.32.0 Released: Native Support for LTX 2.5 and Qwen Image 3.0 — Gremlation · 2026-08-13