Qwen Audio Realtime Tested: Cost and Reasoning vs GPT Models

ArtificialAnlys · x · 2026-07-29

Artificial Analysis has released a new comparison of native Speech-to-Speech models. On the Big Bench Audio subset, Qwen Audio 3.0 Realtime Plus costs $4.42 per hour of input audio—slightly more expensive than GPT-Realtime-2 High ($4.14) but 2.4x cheaper than GPT-Realtime-2.1 High ($10.75).

Meanwhile, the Flash variant measures $4.77 per hour, appearing more costly than Plus on the benchmark. This is driven by comparatively more verbose responses: Flash generates an average of 2,043 total output tokens per response versus 1,232 for Plus.

Related event: Qwen Audio 3.0 Realtime Costs Slightly Higher Than GPT in Benchmark(2 posts)→

Original post →

More from Models

Models channel →