SlopCodeBench: Measuring how sloppy LLM-generated code really is

mitsuhiko · x · 2026-09-11

Sebastian at Earendil digs into SlopCodeBench, tackling the question 'if coding is solved, what now?': LLMs generate formally correct code, but it often introduces unnecessary abstractions, duplicates, and bad decisions. Projects now add millions of LOC per month, eroding human agency—and agents can't clean up the slop either. He finds the industry's approach to measuring code quality remains largely vibes-based.

Original post →

More from coding & agent

coding & agent channel →