Qwen3.8-Omni-Flash cuts overlapping-speech error rate from 88% to 3% and drops audio API pricing 98%

karminski3 · x · 2026-09-18

Qwen released Qwen3.8-Omni-Flash alongside two frameworks, Qwen-Live-Harness and Qwen-MM-Plugins. Key points: +25% average score over Qwen3.5-Omni-Flash with agent capabilities up to 2x; overlapping multi-speaker speech recognition error rate down from 88% to 3%; native support for up to 1 hour of continuous audio/video input; and a 98% API price cut — audio input drops from 18 to 0.8 yuan per million tokens. A ultra-low-latency Qwen3.8-Omni-Flash-Realtime was also released. Caveats: Qwen-Live-Harness has no code pushed to GitHub yet (npm package exists) and currently supports macOS only.

Related event: Qwen3.8-Omni-Flash Launches: Multi-Speaker Error Rate Drops to 3%, API Price Cut 98%(2 posts)→

Original post →

More from Models

Models channel →