OpenAI Finds 30% of Questions in Popular AI Coding Test Corrupted
The Decoder · rss · 2026-07-09
OpenAI reviewed SWE-Bench Pro, a well-known benchmark for measuring AI coding capabilities, and found that around 30% of its tasks are flawed or broken. Consequently, OpenAI decided to withdraw its previous endorsement and support for this benchmark.
More from Research
- OpenAI-style autonomous researchers could become real scientific collaborators — Promptmethus · 2026-07-21
- Soft Clamp cuts tool-call overuse in multi-teacher distillation, from 13.7% to 9.0% — antgroup · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- A developer maps out six design rules for CLIs that humans and AI agents can both use — yujiezha · 2026-07-21
- GPT 5.6 vs. Claude Fable tested in Dyad AI for Physical AI model tuning — ChrisRackauckas · 2026-07-21