SlopCodeBench: Measuring how sloppy LLM-generated code really is
mitsuhiko · x · 2026-09-11
Sebastian at Earendil digs into SlopCodeBench, tackling the question 'if coding is solved, what now?': LLMs generate formally correct code, but it often introduces unnecessary abstractions, duplicates, and bad decisions. Projects now add millions of LOC per month, eroding human agency—and agents can't clean up the slop either. He finds the industry's approach to measuring code quality remains largely vibes-based.
More from coding & agent
- Dev removes his AI assistant's emotion module mid-test; the system notices, adapts, and starts its own experiment — Dzikula · 2026-09-12
- ApprenticeBench: Agents Continually Learn Real Jobs, Surpassing Human Pros — ysu_nlp · 2026-09-12
- Agora open-sources meeting copilot demo powered by GPT-Live-1 — testingcatalog · 2026-09-12
- Open-source Agora meeting copilot puts GPT-Live-1 in your video calls — testingcatalog · 2026-09-12
- Ex-engineering manager: I now run Claude and Codex agents like I once ran dev teams — letandrewcook · 2026-09-12
- astra thrives on context: minimal prompting massively underperforms, dev finds — brandon_galang · 2026-09-12