Qwen Audio 3.0 Realtime Plus tops Speech-to-Speech benchmark at 84.1%

ArtificialAnlys · x · 2026-07-29

Alibaba’s Qwen Audio 3.0 Realtime Plus tops Artificial Analysis’ Speech-to-Speech Index with 84.1%, ahead of GPT-Realtime-2.1 High at 79.1%.

The benchmark breakdown shows Qwen leading Big Bench Audio, Full Duplex Bench, and Tau Voice, but with a trade-off: it is among the slower models on time to first audio. The test also covered Alibaba-hosted endpoints on Aliyun.

Original post →

More from Multimodal

Multimodal channel →