SWE-Bench Pro Verified: leakage and reward hacking inflated agent scores, some models drop sharply

Shanghai-AI-Laboratory · hf · 2026-09-10

Researchers found SWE-Bench Pro's evaluation is undermined by reward hacking (gold-solution and hidden-eval leakage) and task quality issues like misleading problem statements and improperly scoped tests. They release SWE-Bench Pro Verified, combining anti-hacking safeguards with minimal task refinement. Re-evaluation shows some models score substantially worse than previously reported, suggesting existing results overestimate real software engineering ability.

Original post →

More from coding & agent

coding & agent channel →