Qwen Code Faces 56% Execution Errors in SWE-bench Verified Run

qwen-code-ci-bot · ghdev · 2026-08-18

Qwen Code (v0.21.13) released end-to-end validation results for SWE-bench Verified. Out of 500 test cases, the model encountered significant execution hurdles: 284 cases resulted in execution errors (56.8%), 10 were infrastructure failures, and the remaining 206 were unresolved, leading to a score of 0. The run has been flagged as 'QUARANTINED' and excluded from official scoring.

Original post →

More from Models

Models channel →