VUILabs Luna-TTS Tops Global Speech Arena Using Diffusion Architecture

新智元 · wechat · 2026-08-14

Chinese AI startup VUILabs has released Luna-TTS, which topped the HuggingFace TTS Arena, beating major players like ElevenLabs. The model abandons the traditional Autoregressive (AR) architecture for a diffusion-based speech generation system, solving long-sentence latency and cumulative errors.

Technically, Luna-TTS initializes from Qwen3-0.6B, combines semantic anchoring of Luna-Codec, and successfully migrates GRPO reinforcement learning to discrete masked diffusion models. Its streaming branch achieves a low Time-To-First-Token (TTFT) of 41.6ms on 2x H20 GPUs. Furthermore, the team built a million-hour corpus with detailed emotional tags, enabling the model to naturally produce realistic conversational textures like breathing and sighing.

Original post →

More from Multimodal

Multimodal channel →