VUILabs Luna-TTS Tops Global Speech Arena Using Diffusion Architecture
新智元 · wechat · 2026-08-14
Chinese AI startup VUILabs has released Luna-TTS, which topped the HuggingFace TTS Arena, beating major players like ElevenLabs. The model abandons the traditional Autoregressive (AR) architecture for a diffusion-based speech generation system, solving long-sentence latency and cumulative errors.
Technically, Luna-TTS initializes from Qwen3-0.6B, combines semantic anchoring of Luna-Codec, and successfully migrates GRPO reinforcement learning to discrete masked diffusion models. Its streaming branch achieves a low Time-To-First-Token (TTFT) of 41.6ms on 2x H20 GPUs. Furthermore, the team built a million-hour corpus with detailed emotional tags, enabling the model to naturally produce realistic conversational textures like breathing and sighing.
More from Multimodal
- 18-sec One-Shot! Seedance 2.5 Maldives Resort Video Prompt — techhalla · 2026-08-14
- DotSight AI Releases Open-Source 280B Multimodal Model with Apache 2.0 License — Xianbao_QIAN · 2026-08-14
- Wan 3.0 official prompts revealed: 64 prompts as 30-second shot scripts — Tricky_Algae2625 · 2026-08-14
- AI horrorcore diss track 'Syko Sam' released, showcasing AI music creation — JonMillaTheKilla · 2026-08-14
- Seedance2.5 Generates One-Sentence Rock Music Video, Blurring Reality and AI — Eric520CC · 2026-08-14
- MiniMax i2v Workflow: Combining fl2va and ref2va for Voice Cloning — spiderofmars · 2026-08-14