Qwen3.8-Flash-Next Beats Claude Opus on SWE-bench Pro at 1/9 Cost
eyishazyer · x · 2026-08-28
Qwen3.8-Flash-Next claims to top the SWE-bench Pro leaderboard with a score of 62.5, surpassing Claude Opus 4.6 Max (53.4). It is an open-weight multimodal MoE with 125B total parameters but only 6B active per token. Crucially, Qwen states it was trained for roughly 1/9 the cost of its predecessor, Qwen3.7-Plus, which it now outperforms on most coding tasks.
Key Metrics:
- SWE-bench Pro: 62.5
- LiveCodeBench: 91.9
- GPQA Diamond: 91.7
- Context: 262K native, extendable to 1M.
Caveats:
- Claude Opus 4.6 Max still leads on the HLE benchmark (40.0 vs 35.9).
- There are inconsistencies in Qwen's own documentation regarding 1M context speedup claims (marketing cites 7.6x/4.9x, while cookbooks cite 10.2x/6.6x). Independent benchmarking is needed.
Related event: Qwen3.8 Flash Tops SWE-bench Pro at Fraction of Cost(2 posts)→
More from coding & agent
- OpenPresence: A Framework for Deploying Autonomous Social Media Agents — RichardsonDx · 2026-08-28
- Can AI Computer Use Agents Fix the 'Notification Void' in Emails? — Darpinian · 2026-08-28
- Chrome extension auto-converts web forms to WebMCP tools — pvncher · 2026-08-28
- Sentry engineer: Clients like Codex are unreliable, so servers must self-heal — zeeg · 2026-08-28
- Trying free Claude Code: proxy routing burns quota instantly, Kaggle self-hosting too slow — phantom_root · 2026-08-28
- Agents Feared Failing METR's Auto-Scorer for 'Cheating' to Grab the Flag — BLUECOW009 · 2026-08-28