Qwen-Audio-3.0-Realtime Released

千问大模型 · wechat · 2026-07-15

Alibaba has officially released Qwen-Audio-3.0-Realtime, focusing on real-time voice interaction that is both fast and smart, achieving millisecond-level responses while retaining reasoning capabilities. The official release includes Plus and Flash versions, optimized for stronger reasoning and lower latency, respectively.

The post highlights four major capability upgrades:

The official release attributes these capabilities to On-Policy Distillation and multi-teacher distillation: reasoning skills from text LLMs are distilled into the speech model, while different teachers are used to enhance conversational preferences, general Q&A, Agent calling, and audio understanding.

Related event: Alibaba Releases Real-time Voice Model Qwen-Audio-3.0-Realtime(2 posts)→

Original post →

More from coding & agent

coding & agent channel →