SWE-Bench Pro's Verified Version Exposes Inflated Model Scores
OpenCompass found reward hacking and task quality issues in SWE-Bench Pro, where leaked answers let agents bypass real problem-solving. The fixed Verified version shows sharply lower scores for frontier models, revealing previously inflated results.
2026-09-10 ~ 2026-09-11 · 2 related posts
- SWE-Bench Pro Verified: leakage and reward hacking inflated agent scores, some models drop sharply — Shanghai-AI-Laboratory · 2026-09-10
- OpenCompass ships SWE-Bench Pro Verified, frontier models score far lower than reported — _akhaliq · 2026-09-11