Qwen Team Details Qwen3.8-Omni-Flash: an Omni-Modal Agent Model with 1M Context
dair_ai · x · 2026-09-26
The Qwen team published a report on their omni-modal agents, introducing Qwen3.8-Omni-Flash, a natively multimodal model trained for long-horizon agent tasks across text, audio, and video — e.g., video editing and long-form audio/video translation.
Key points:
- Built on the sparse MoE design of Qwen3.8-Next, with a 1M-token context window;
- A co-training strategy preserves text performance while transferring agent skills to audio/video tasks;
- Ships with two open-source frameworks: Qwen-MM-Plugins adds audio/video support to existing agent harnesses, and Qwen-Live-Harness handles real-time multimodal interaction with context/memory management, tool use, and sub-agent delegation;
- A paper accompanies the release.
Related event: Qwen Releases Omni-Modal Agent Model Qwen3.8-Omni-Flash(2 posts)→
More from coding & agent
- Armin Ronacher: with agents this good, Rust devs should stop fearing unsafe — mitsuhiko · 2026-09-26
- User replaces Claude with local Qwen 3.8 Next on a 128GB Mac for days — AdInternational5848 · 2026-09-26
- Agent passports? Reddit debates how websites should admit good bots — the-real-news · 2026-09-26
- Casual DSPy Webinar on Jev Support and Design Patterns in the Works — dbreunig · 2026-09-26
- How the 134KB realtime city demo was made: Opus builds its own tiny asset tools — gandamu_ml · 2026-09-26
- Xcode 27 exposes 50+ tools via MCP for coding agents like Claude and Codex — JordanMorgan10 · 2026-09-26