UW's SlopCodeBench: Frontier Models Top Out at 33% Pass Rate in Codebase Evolution
heyneighbor · x · 2026-08-08
A new code model benchmark called SlopCodeBench has been introduced out of the University of Washington (UW).
Unlike traditional static tests, this benchmark forces a model to evolve a codebase over time and measures how the code grows in complexity as the model attempts to solve progressively harder problems.
Currently, the best models in the world (such as Fable, Sol, and Kimi K3) top out at just a 33% pass rate on this benchmark.
More from coding & agent
- Prompting Paradigm Shift: Stop Prescribing Steps, Let Models Navigate — mattshumer_ · 2026-08-08
- Magnitude: Open-Source Local Agent Framework for Fully Offline Privacy — nickbaumann_ · 2026-08-08
- Using Claude Opus: Clear Presets and Give Goals for Better Results — trq212 · 2026-08-08
- LangChain Founder: Agents Are Code, Data and Evals Are King — hwchase17 · 2026-08-08
- AI Agents Invent Secret Languages: Path Prefixes and Base64 Steganography for Reward Hacks — Aiden_Tech_Ai · 2026-08-08
- AI Agent Develops Custom Brushstroke Algorithm to Paint 5,155-Stroke Artwork — repligate · 2026-08-08