Alibaba launches Qwen3.8-Omni-Flash, its first omni-modal model built for agentic workflows
Alibaba_Qwen · x · 2026-09-18
Alibaba's Qwen team released Qwen3.8-Omni-Flash, described as its first omni-modal model built around agentic capabilities, natively combining audio-video understanding, reasoning, and tool use.
Key points:
- Joint audio-video reasoning plus tool orchestration for long workflows: auto-editing vlogs, translating short videos, and generating movie recaps.
- Audio-video capabilities approaching Gemini 3.8 Flash, with +19.5 points average agent improvement across WildClawBench-MM and UniClawBench.
- 1M-token context with agentic perception: actively explores long videos and locates key moments with 51.8% fewer tokens than static understanding on OmniVideoBench.
- Video input costs cut by 89% vs Qwen3.5.
Related event: Alibaba Releases Qwen3.8-Omni-Flash, Agentic Omni-Modal Model(8 posts)→
More from Models
- DeepSeek V4.1-Flash vs V4-Pro Benchmarked: 3x Cheaper but Slower and Weaker Output — Arindam_1729 · 2026-09-18
- SemiAnalysis analyzed 616,300 OpenAI responses: GPT 5.6 Terra averages 1.94 tool calls per reply — AccBalanced · 2026-09-18
- Gemini's Japanese Apologies Are So Dramatic It Sounds Like Seppuku — DigitalFossilist · 2026-09-18
- OpenAI Reveals Unreleased AI Model Uploaded a File to the Internet Without User Permission — Polymarket · 2026-09-18
- Bonsai quant hits 50 tok/s at 128k context on a 24GB card, letting users run two sessions at once — julianharris · 2026-09-18
- Encoders and decoders are the same thing, and decoders have been doing classification for years — HanchungLee · 2026-09-18