Audit of SWE Verified suggests 5%–10% of samples may still be invalid

xeophon · x · 2026-07-25

A year-old audit of SWE Verified is used to argue that the benchmark still contains a meaningful number of bad samples. The post says a careful inspection of the top five SWE-bench submissions found remaining problematic cases, including unclear issue descriptions and overly restrictive tests.

The analysis claims:

The point is that even a widely used coding benchmark can still carry enough noise to affect leaderboard interpretation.

Related event: SWE-bench Audit Reveals Up to 10% Invalid Samples(2 posts)→

Original post →

More from Research

Research channel →