Qwen Code Scores 76.16% on Full SWE-bench Verified Run, Resolving 377 of 500
DennisYu07 · ghdev · 2026-08-21
QwenLM/qwen-code posted end-to-end full benchmark results (v0.21.15, model qwen3.7-plus):
- SWE-bench Verified: all 500 tasks completed, 377 resolved, 118 unresolved, 2 execution errors, 3 infrastructure failures, scoring 76.16% (denominator: 495 valid grader results).
- Followed by Terminal-Bench 2.0 (89 tasks) with verifier-backed results and trajectory writeback.
- Case trajectories are packaged as swe-bench-verified-dsw-eas-full-20260821-r1-trajectories.tar.gz; the full run is available on GitHub Actions.
More from coding & agent
- Agentic coding accessibility will reshape understanding of software complexity — pixlpa · 2026-08-24
- Devin Agent bypasses Slack block by finding emails in git logs — sandylikesfrogs · 2026-08-24
- Developer habits shift: Agents become collaborators from simple tools — latticecut · 2026-08-24
- Dev bottleneck shifts from writing to reading code: exe.dev co-founder — thursdai_pod · 2026-08-24
- The biggest AI mistake: trying to reinvent the wheel instead of using tools — Tired40s · 2026-08-24
- DeepPaperNote turns research papers into Obsidian notes — tom_doerr · 2026-08-24