Qwen-Audio-3.0-Realtime Upgrade
通义实验室 · wechat · 2026-07-15
Tongyi Lab has released an upgraded real-time voice interaction model, **Fun-Realtime-AudioChat 升级版**, now renamed **Qwen-Audio-3.0-Realtime**. The update focuses on making voice conversations more natural and human-like: - **情感与语气表达**: Dynamically adjusts tone, rhythm, pitch, and emotion based on context, featuring voice cloning to reduce the robotic feel. - **多模态双工控制**: More stable identification of the primary speaker in noisy environments, resisting interruptions from background noise, interjections, or multi-person chats. Offers voiceprint-level background filtering. - **工具调用能力**: Automatically determines when to invoke tools based on context, supporting route planning, info queries, scheduling, and table lookups, emphasizing multitasking while chatting. - **底座与评测**: Uses On-Policy Distillation and multi-teacher distillation to transfer reasoning, agentic, and audio understanding capabilities from text models to the voice model. The post shares benchmark scores from VoiceBench, AudioMultiChallenge, and ArtificialAnalysis, claiming SOTA or overall first place in several sub-items. The model is now available on Alibaba Cloud Bailian, with links to APIs, documentation, and demos provided.
Related event: Alibaba Releases Real-time Voice Model Qwen-Audio-3.0-Realtime(2 posts)→
More from coding & agent
- Why vector databases slow AI agents down after constant writes — PrajwalTomar_ · 2026-07-21
- A 13-minute GitHub Copilot video digs into prompt caching — lee_stott · 2026-07-21
- Codex turns out 123 screensavers in one playful batch — intellectronica · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- CHAP defines approvals, handoffs, and audit logs for human-agent workflows — DeliveryTechnical199 · 2026-07-21
- The author says Codex reached 20x and is now debugging spec decoding on a hybrid parallel setup — TheZachMueller · 2026-07-21