Alibaba launches Qwen3.8-Omni-Flash: agent-first omni-modal model with 1M context

johnseach · x · 2026-09-18

Qwen released Qwen3.8-Omni-Flash, its first omni-modal model built around agentic capabilities: native text/image/audio/video input, up to 1M-token context, and tool orchestration for workflows like auto-editing vlogs and video recaps. Qwen claims audio-video performance approaching Gemini 3.8 Flash, with +19.5 points average on agent benchmarks (WildClawBench-MM & UniClawBench). Priced at $0.15/M input tokens, $0.47/M output ($0.016/M with implicit cache); compatible with DashScope and OpenAI protocols, with a Qwen-MM-Plugins companion for agent frameworks.

Related event: Alibaba Launches Qwen3.8-Omni-Flash: Agentic Omni-Modal Model with Over 98% Price Cut(15 posts)→

Original post →

More from coding & agent

coding & agent channel →