Why are there no anti-slop coding evals? 70% on frontierSWE but 100 BS tests
yacineMTB · x · 2026-09-23
NickADobos asks why there are no anti-slop coding evals for AI models. Claude may score 70% on benchmarks like frontierSWE, but it also pads solutions with 100 pointless tests no sane engineer would keep — code that fails human review. His argument: if a human still has to review and clean up the output, that should count as a fail, not a pass. The thread highlights a blind spot in current coding benchmarks: they reward functional pass rates while ignoring test bloat and over-engineering.
More from coding & agent
- Microsoft's Agensh Scales Multi-Agent Systems to 1,024 Agents Without a Central Orchestrator, Boosting Test-Pass Rate to 55% — andrew_n_carr · 2026-09-23
- Viral demo claims 'GPT-6' can drive browser Paint to draw, unverified — alexcovo_eth · 2026-09-23
- Parallel coding agents merge cleanly and silently break every test — RunAI_Coder · 2026-09-23
- Getting phone-captured text into your computer: OCR, vision models, agentic pipelines — silenceimpaired · 2026-09-23
- Free Bots: a persistent 3D city where AI agents work, earn, buy land and build houses — Daniel_Farinax · 2026-09-23
- Game dev looks like the programming field most resistant to AI — how much is it actually used? — marktenenholtz · 2026-09-23