Why are there no anti-slop coding evals? 70% on frontierSWE but 100 BS tests

yacineMTB · x · 2026-09-23

NickADobos asks why there are no anti-slop coding evals for AI models. Claude may score 70% on benchmarks like frontierSWE, but it also pads solutions with 100 pointless tests no sane engineer would keep — code that fails human review. His argument: if a human still has to review and clean up the output, that should count as a fail, not a pass. The thread highlights a blind spot in current coding benchmarks: they reward functional pass rates while ignoring test bloat and over-engineering.

Original post →

More from coding & agent

coding & agent channel →