Qwen3.8-Max fixes 19 of 105 hidden bugs in blind coding benchmark
breath_mirror · x · 2026-08-04
A blind bug-hunt benchmark puts Qwen3.8-Max in the middle of the pack for coding, despite Alibaba’s “new bar for coding” pitch.
- The test covered two real repos, 105 hidden bugs, 11 frontier models, and 15 runs.
- Qwen3.8-Max fixed 19 bugs in 148 minutes at a cost of $31.10.
- By comparison, GPT-5.6 Sol led with 42 fixes, while Kimi K3 and Opus 5 each fixed 21.
- The author says Qwen did find one bug no other model found, but getting the model running took several attempts and some subscription/account friction.
- A takeaway from the benchmark: GPT-5.6 Luna did the same benchmark for $1.80 and fixed 33, while Grok 4.5 fixed 16 in under 25 minutes.
- The benchmark image also notes that 52 of the 105 bugs survived every model.
More from coding & agent
- Claude 3.5 Sonnet Aces Adversarial Data Science Test, Catches Data Leakage Autonomously — hugobowne · 2026-08-04
- GitHub list tracks OSINT MCP servers for Claude, Cursor and Windsurf — tom_doerr · 2026-08-04
- Vibe coding gives old tools like MediaPipe a second life — bilawalsidhu · 2026-08-04
- Google details the pipeline behind its open-source Agent Skills — _jaydeepkarale · 2026-08-04
- Steve Yegge says Opus 4.7 kept tweaking Gas Town instead of finishing work — Simon Willison · 2026-08-04
- OpenClaw warning says agents with web access, chat exposure and local computer control are unsafe — ericelliott_ · 2026-08-04