OpenCompass ships SWE-Bench Pro Verified, frontier models score far lower than reported

_akhaliq · x · 2026-09-11

OpenCompass released a verified version of SWE-Bench Pro that fixes reward hacking and task-quality issues in the original benchmark. The corrected results show frontier models score far lower than previously reported, suggesting systematic inflation in agent coding leaderboard numbers.

Related event: SWE-Bench Pro's Verified Version Exposes Inflated Model Scores(2 posts)→

Original post →

More from Models

Models channel →