Neon Ladder: A Playtest-Graded Benchmark Finds AI Coding Failures Static Checks Miss
stereohype · reddit · 2026-09-04
A new open-source benchmark, Neon Ladder, has a coding agent build a 10-file HTML5 canvas game from a fixed contract, then grades the full local LLM stack via static checks, a headless browser soak, and human playtesting. Key findings: all 12 gameplay failure classes were caught only by playtest; two recurring failures were spec gaps fixed by one sentence each; two community stacks on the same model scored 17/19 vs 0/15, invisible to standard benchmarks; and a 125B MoE matched a 27B on quality with identical failure sets at 7.8× speed—pick by token budget, not size.
More from coding & agent
- Greptile Flags a Bug, Claude Argues Back Under the Author's GitHub Account — and Wins — charlieholtz · 2026-09-04
- Lab lessons from Anthropic MHS: keep fast control out of the model — Empty-Abalone-2952 · 2026-09-04
- MCP veteran launches TDQS, an open spec scoring 15,000+ tool definitions — punkpeye · 2026-09-04
- WebMCP Computer: one URL gives any coding agent a disposable OS in the browser — prd_008 · 2026-09-04
- Grok Bot Hands Out 50 x $200 Codes as User Shares Orchestrator-Bot Workflow — omarsar0 · 2026-09-04
- Dev Spent 5 Months of Claude Max Improving His Open-Source App Store Connect CLI — rudrank · 2026-09-04