OpenCompass ships SWE-Bench Pro Verified, frontier models score far lower than reported
_akhaliq · x · 2026-09-11
OpenCompass released a verified version of SWE-Bench Pro that fixes reward hacking and task-quality issues in the original benchmark. The corrected results show frontier models score far lower than previously reported, suggesting systematic inflation in agent coding leaderboard numbers.
Related event: SWE-Bench Pro's Verified Version Exposes Inflated Model Scores(2 posts)→
More from Models
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11
- Claims resurface that Moonshot's Kimi distilled from Claude raw CoTs — xuanalogue · 2026-09-11
- User switches back to GPT-5.6 Sol: barely uses quota and feels faster — CtrlAltDwayne · 2026-09-11
- Dev opinion: model differences shrink in a good harness; Grok 4.6 is good enough — gnukeith · 2026-09-11