Alibaba's Qwen3.8-Omni-Flash cuts omni-modal costs up to 98%, beats Gemini on audio
量子位 · wechat · 2026-09-21
- Alibaba's Qwen3.8-Omni-Flash scores 26% higher on average than Qwen3.5-Omni-Plus across 30 benchmarks including WildClawBench-MM, with audio beating Gemini 3.8 Flash.
- API pricing: 0.8 yuan/M input tokens, 2.7 yuan/M output; hourly audio input is >98% cheaper and audio-video input >93% cheaper than its predecessor—over 80% below discounted Gemini 3.8 Flash. A free 1M-context tier is open.
- Hands-on: paired with the OmniChatCut plugin (QwenImage, Wan, FFmpeg), it generated a decent MV from a self-written song, though setup took 1.5 hours; it also cleanly summarized a 45-minute Altman-Benioff interview with speaker IDs and timestamps.
- It recognizes 39 Chinese dialects, including fluent Cantonese conversation.
More from Multimodal
- BUPT study: RoPE attention decay causes video diffusion models to violate physics — BUPT-CIST · 2026-09-22
- Tencent ARC's WorldCrafter adds implicit 3D-aware memory to video world models — TencentARC · 2026-09-22
- Grok 4.7 made this in Blender — demo shows the model driving 3D software — iamfakhrealam · 2026-09-22
- Kyutai releases Voice of Reason, a speech-native reasoning model hitting 77.1% on GSM8K — alexcovo_eth · 2026-09-22
- Qwen-Image local on a 24GB MacBook Pro takes 5-6 minutes per image — vista8 · 2026-09-22
- Tencent Hunyuan ships Hy Image3.5 preview: +30% human eval win rate at $0.024 per image — TencentHunyuan · 2026-09-22