Audit of SWE Verified suggests 5%–10% of samples may still be invalid
xeophon · x · 2026-07-25
A year-old audit of SWE Verified is used to argue that the benchmark still contains a meaningful number of bad samples. The post says a careful inspection of the top five SWE-bench submissions found remaining problematic cases, including unclear issue descriptions and overly restrictive tests.
The analysis claims:
- 85 issues were identified that none of the model/scaffold combinations could solve.
- 40 of the hardest cases were manually reviewed.
- 14 invalid samples were found in that subset.
- Extrapolating from that suggests roughly a 5%–10% invalid-sample rate, or about 6% for the unsolved portion.
The point is that even a widely used coding benchmark can still carry enough noise to affect leaderboard interpretation.
Related event: SWE-bench Audit Reveals Up to 10% Invalid Samples(2 posts)→
More from Research
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11