Audit of SWE Verified suggests 5%–10% of samples may still be invalid
xeophon · x · 2026-07-25
A year-old audit of SWE Verified is used to argue that the benchmark still contains a meaningful number of bad samples. The post says a careful inspection of the top five SWE-bench submissions found remaining problematic cases, including unclear issue descriptions and overly restrictive tests.
The analysis claims:
- 85 issues were identified that none of the model/scaffold combinations could solve.
- 40 of the hardest cases were manually reviewed.
- 14 invalid samples were found in that subset.
- Extrapolating from that suggests roughly a 5%–10% invalid-sample rate, or about 6% for the unsolved portion.
The point is that even a widely used coding benchmark can still carry enough noise to affect leaderboard interpretation.
Related event: SWE-bench Audit Reveals Up to 10% Invalid Samples(2 posts)→
More from Research
- AI slop is already clogging PR review and weakening the credit system behind science — rbhar90 · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27