SWE-Bench Pro Audit Reveals Issues
dl_weekly · x · 2026-07-13
An OpenAI audit discovered that about 30% of the tasks in SWE-Bench Pro are flawed, prompting the withdrawal of its previous recommendation as an alternative to SWE-bench Verified.
This indicates that even highly anticipated benchmarks for software engineering can contain defects severe enough to compromise their reliability. For those using benchmarks for model selection, comparison, or evaluating coding agents, such "benchmark audits" serve as crucial warning signs.
More from coding & agent
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11
- How Do You Catch Behavioral Regressions in LLM Agents Between Releases? — Beautiful_Belt_601 · 2026-09-11
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- Running the Firefox MCP on Android via Termux, ngrok, and mcp-proxy — Nervous-Strain7544 · 2026-09-11
- Run Firefox MCP on Android: Termux + ngrok tunnel tutorial — Nervous-Strain7544 · 2026-09-11