SWE-Bench Pro Audit Reveals Issues
dl_weekly · x · 2026-07-13
An OpenAI audit discovered that about 30% of the tasks in SWE-Bench Pro are flawed, prompting the withdrawal of its previous recommendation as an alternative to SWE-bench Verified.
This indicates that even highly anticipated benchmarks for software engineering can contain defects severe enough to compromise their reliability. For those using benchmarks for model selection, comparison, or evaluating coding agents, such "benchmark audits" serve as crucial warning signs.
More from coding & agent
- Dev builds interactive 3D product experience with GPT-6 Astra + Hyper3D Rodin — nikola_mr64990 · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11