Alibaba launches Qwen3.8-Omni-Flash: agent-first omni-modal model with 1M context
johnseach · x · 2026-09-18
Qwen released Qwen3.8-Omni-Flash, its first omni-modal model built around agentic capabilities: native text/image/audio/video input, up to 1M-token context, and tool orchestration for workflows like auto-editing vlogs and video recaps. Qwen claims audio-video performance approaching Gemini 3.8 Flash, with +19.5 points average on agent benchmarks (WildClawBench-MM & UniClawBench). Priced at $0.15/M input tokens, $0.47/M output ($0.016/M with implicit cache); compatible with DashScope and OpenAI protocols, with a Qwen-MM-Plugins companion for agent frameworks.
More from coding & agent
- Devs debate running stateful AI agent runtimes on Cloudflare Workers and other edge runtimes — merlinofthewater · 2026-09-20
- Jev's instant compression scores each tool call to trim agent context without summarization — FinanceYF5 · 2026-09-20
- Stop using LLMs for ticket triage: Jev turns unstructured input into an executable smart if — sven_ai · 2026-09-20
- Jeff Dean's 1-Hour AI Engineering Lecture: From LLM Basics to Agent Graphs — irinarish · 2026-09-20
- Multi-agent systems work best with clear roles, not more agents — _jaydeepkarale · 2026-09-20
- Jev founder: all AI models are built for human-in-the-loop, not true automation — hardimanjames · 2026-09-20