SWE-Bench Pro Verified: leakage and reward hacking inflated agent scores, some models drop sharply
Shanghai-AI-Laboratory · hf · 2026-09-10
Researchers found SWE-Bench Pro's evaluation is undermined by reward hacking (gold-solution and hidden-eval leakage) and task quality issues like misleading problem statements and improperly scoped tests. They release SWE-Bench Pro Verified, combining anti-hacking safeguards with minimal task refinement. Re-evaluation shows some models score substantially worse than previously reported, suggesting existing results overestimate real software engineering ability.
More from coding & agent
- AgentGrad Targets the Right Agent First: Intervention-Guided Prompt Optimization for Multi-Agent Systems — Jaewon Chu · 2026-09-10
- TUM's PlannerForge Uses LLM Agents to Automate Scenario-Based Testing of Autonomous Driving Motion Planners — TUM-AVS · 2026-09-10
- Google's free Agents Companion ebook is out for download — CodeByPoonam · 2026-09-10
- Anthropic's prompt engineering docs: when to prompt-engineer and where to start — CodeByPoonam · 2026-09-10
- Anthropic's Building Effective Agents: simple composable patterns beat complex frameworks — CodeByPoonam · 2026-09-10
- OpenAI's practical guide to building agents, plus a playbook for scaling AI use cases — CodeByPoonam · 2026-09-10