UW's SlopCodeBench: Frontier Models Top Out at 33% Pass Rate in Codebase Evolution

heyneighbor · x · 2026-08-08

A new code model benchmark called SlopCodeBench has been introduced out of the University of Washington (UW).

Unlike traditional static tests, this benchmark forces a model to evolve a codebase over time and measures how the code grows in complexity as the model attempts to solve progressively harder problems.

Currently, the best models in the world (such as Fable, Sol, and Kimi K3) top out at just a 33% pass rate on this benchmark.

Original post →

More from coding & agent

coding & agent channel →