Qwen3.8-Omni-Flash cuts overlapping-speech error rate from 88% to 3% and drops audio API pricing 98%
karminski3 · x · 2026-09-18
Qwen released Qwen3.8-Omni-Flash alongside two frameworks, Qwen-Live-Harness and Qwen-MM-Plugins. Key points: +25% average score over Qwen3.5-Omni-Flash with agent capabilities up to 2x; overlapping multi-speaker speech recognition error rate down from 88% to 3%; native support for up to 1 hour of continuous audio/video input; and a 98% API price cut — audio input drops from 18 to 0.8 yuan per million tokens. A ultra-low-latency Qwen3.8-Omni-Flash-Realtime was also released. Caveats: Qwen-Live-Harness has no code pushed to GitHub yet (npm package exists) and currently supports macOS only.
More from Models
- ChatGPT co-inventor launches Jev, claiming 200x faster, 400x cheaper frontier model — multiply_matrix · 2026-09-18
- Tencent's Hy4 Preview ranks #4 among open-weight models, cheapest in top ten — mariofilhoml · 2026-09-18
- Simple letter-counting test exposes huge gap: GPT-6-Astra hits 93%, Fable 5.1 flounders — scaling01 · 2026-09-18
- Jev, a 'System One' model by Typesafe, launches on OpenRouter with typed decisions instead of text — majidmanzarpour · 2026-09-18
- Zhipu claims AI autonomously discovered a WeWorm-exploitable vulnerability — teortaxesTex · 2026-09-18
- User comparison: Gemini nailed an insurance-law question that ChatGPT defended with circular reasoning — Hatrct · 2026-09-18