OpenAI Says SWE-Bench Pro Is Too Noisy
OpenAI says about 30% of tasks in the coding benchmark SWE-Bench Pro are problematic, weakening its reliability as a measure of frontier programming ability. It highlighted four major issues: overly strict tests, unclear prompts, low test coverage, and misleading prompt wording.
2026-07-09 ~ 2026-07-10 · 3 related posts
- OpenAI Questions SWE-Bench Pro Evaluation — FateOfMuffins · 2026-07-09
- OpenAI Points Out Flaws in SWE-Bench Pro — rseroter · 2026-07-10
1 near-duplicate retellings: RishiBommasani