OpenAI Finds 30% of Questions in Popular AI Coding Test Corrupted

The Decoder · rss · 2026-07-09

OpenAI reviewed SWE-Bench Pro, a well-known benchmark for measuring AI coding capabilities, and found that around 30% of its tasks are flawed or broken. Consequently, OpenAI decided to withdraw its previous endorsement and support for this benchmark.

Original post →

More from Research

Research channel →