OpenAI Says SWE-Bench Pro Is Too Noisy

OpenAI says about 30% of tasks in the coding benchmark SWE-Bench Pro are problematic, weakening its reliability as a measure of frontier programming ability. It highlighted four major issues: overly strict tests, unclear prompts, low test coverage, and misleading prompt wording.

2026-07-09 ~ 2026-07-10 · 3 related posts

1 near-duplicate retellings: RishiBommasani