Qwen Audio 3.0 Realtime Plus tops Speech-to-Speech benchmark at 84.1%
ArtificialAnlys · x · 2026-07-29
Alibaba’s Qwen Audio 3.0 Realtime Plus tops Artificial Analysis’ Speech-to-Speech Index with 84.1%, ahead of GPT-Realtime-2.1 High at 79.1%.
The benchmark breakdown shows Qwen leading Big Bench Audio, Full Duplex Bench, and Tau Voice, but with a trade-off: it is among the slower models on time to first audio. The test also covered Alibaba-hosted endpoints on Aliyun.
More from Multimodal
- A local iPhone photo editor enters beta with on-device vision models — measure_plan · 2026-07-29
- Open-source app uses Gemini Video Understanding to pull highlight reels automatically — icnahom · 2026-07-29
- Researchers want agent runs to end with short explainer videos, not text walls — airesearch12 · 2026-07-29
- WAN Bernini plus Prompt Relay gives finer control over 10–15 second videos — Sudden_List_2693 · 2026-07-29
- HeyGen schedules a HyperFrames live demo at Seattle Tech Week on July 29 — toolstelegraph · 2026-07-29
- HeyGen Open-Sources HyperFrames: Natural Language Video Grading & Rendering — toolstelegraph · 2026-07-29