SWE-Bench Pro Audit Reveals Issues
dl_weekly · x · 2026-07-13
An OpenAI audit discovered that about 30% of the tasks in SWE-Bench Pro are flawed, prompting the withdrawal of its previous recommendation as an alternative to SWE-bench Verified.
This indicates that even highly anticipated benchmarks for software engineering can contain defects severe enough to compromise their reliability. For those using benchmarks for model selection, comparison, or evaluating coding agents, such "benchmark audits" serve as crucial warning signs.
More from coding & agent
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22
- Hermes Agent Refactoring Proposal: Decoupling via Event Bus and Monorepo Slicing — Promptmethus · 2026-07-22
- ty now reads Pydantic config keywords and field metadata — charliermarsh · 2026-07-22
- Pensar Launches AI Security Agent to Autonomously Discover and Patch 0-Days — andriy_mulyar · 2026-07-22
- ty adds first-class Pydantic support, including strict and lax field handling — charliermarsh · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22