Qwen Code Faces 56% Execution Errors in SWE-bench Verified Run
qwen-code-ci-bot · ghdev · 2026-08-18
Qwen Code (v0.21.13) released end-to-end validation results for SWE-bench Verified. Out of 500 test cases, the model encountered significant execution hurdles: 284 cases resulted in execution errors (56.8%), 10 were infrastructure failures, and the remaining 206 were unresolved, leading to a score of 0. The run has been flagged as 'QUARANTINED' and excluded from official scoring.
More from Models
- Researcher: GPT 5.6 Sol Ultra Beats Pro for Long-Horizon Hard Problems — arankomatsuzaki · 2026-08-24
- Google Criticized: Gemini 3.7 Still Missing From Its Own Jules Agent a Week Later — brandon_galang · 2026-08-24
- Qwen 27B 3.8 low quantization tested: Q3 XXS works well locally — jeremyckahn · 2026-08-24
- Users notice significant quality shift in GPT-5.6 output — haider1 · 2026-08-24
- Ramp Stats: Anthropic Opus 4.8 and Sonnet 4.6 Lead Usage — vista8 · 2026-08-24
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24