Neon Ladder: A Playtest-Graded Benchmark Finds AI Coding Failures Static Checks Miss

stereohype · reddit · 2026-09-04

A new open-source benchmark, Neon Ladder, has a coding agent build a 10-file HTML5 canvas game from a fixed contract, then grades the full local LLM stack via static checks, a headless browser soak, and human playtesting. Key findings: all 12 gameplay failure classes were caught only by playtest; two recurring failures were spec gaps fixed by one sentence each; two community stacks on the same model scored 17/19 vs 0/15, invisible to standard benchmarks; and a 125B MoE matched a 27B on quality with identical failure sets at 7.8× speed—pick by token budget, not size.

Original post →

More from coding & agent

coding & agent channel →