OpenAI Points Out Flaws in SWE-Bench Pro

rseroter · x · 2026-07-10

OpenAI evaluated SWE-Bench Pro and identified four categories of issues in its code evaluations: overly strict testing, ambiguous prompts, low test coverage, and misleading prompts that skew results.

Related event: OpenAI Says SWE-Bench Pro Is Too Noisy(3 posts)→

Original post →

More from Models

Models channel →