Qwen3.8-Omni-Flash: natively multimodal agent model with 1M-token context
dair_ai · x · 2026-09-26
dairai highlights Qwen team's report on their omni model, calling it the start of "the era of omni models."
- Qwen3.8-Omni-Flash: a natively multimodal model trained for long-horizon agent tasks across text, audio and video (video editing, long-form AV translation).
- Built on Qwen3.8-Next's sparse MoE design with a 1M-token context window; a co-training strategy preserves text performance while transferring agent skills to audio/video.
- Two open-source frameworks ship alongside: Qwen-MM-Plugins (adds AV support to existing agent harnesses) and Qwen-Live-Harness (real-time multimodal interaction with context/memory management, tool use, sub-agent delegation).
- Paper linked in the original post.
More from coding & agent
- Agent Arena Pareto frontier: DeepSeek V4.1 Flash delivers +4.1% at just $0.06 per task — arena · 2026-09-26
- GPT-6 Sol hits #6 on Agent Arena with +7.7% net improvement at 56% less cost per task — arena · 2026-09-26
- "If the app works, the code's perfect": dev says AI ended his code hygiene — banteg · 2026-09-26
- TypeSafe launches Jev, a decision-making model that outputs executable decisions instead of text — hwchase17 · 2026-09-26
- New trick encodes source map data in background colors so AI agents can locate code lines from screenshots — jasonkneen · 2026-09-26
- Agent Arena leaderboard: Claude Fable 5.1 tops with +13.8% net improvement across 2M sessions — arena · 2026-09-26