OpenAI Questions Reliability of SWE-Bench Pro

RishiBommasani · x · 2026-07-10

The post discusses OpenAI's audit of SWE-Bench Pro, revealing that the popular AI coding benchmark can no longer reliably measure cutting-edge programming capabilities. OpenAI identified flaws in about 30% of the questions and subsequently withdrew its previous recommendation of the benchmark as a leading coding evaluation.

Related event: OpenAI Says SWE-Bench Pro Is Too Noisy(3 posts)→

Original post →

More from coding & agent

coding & agent channel →