SlopCodeBench draws praise as a new benchmark for multi-turn code degradation
cedric_chee · x · 2026-09-13
Developer Cedric Chee reviewed the SlopCodeBench docs, problem sets, and repo and called it a well-designed eval, noting that designing good evals is very hard.
- Context: Reports claim GPT-6 Astra's generated code gradually bloats and degrades across turns, with some clever code-golfing behavior — precisely the failure mode this benchmark targets.
- Content: Problems are split into checkpoints (e.g., a medium Python task to build an interactive encrypted password manager), span Python/JavaScript/Java/C++/Go, and are organized by difficulty and category, enabling comparisons of GPT-6 Astra, Fable 5.1, GPT-5.6 Sol, and GLM-5.3 on long-horizon coding tasks.
More from coding & agent
- Developer fixes bug via Grok on iPhone, agent opens PR in minutes — nima_owji · 2026-09-13
- User completes 100% of a week's online shopping via Muse voice agent — armand_ruiz · 2026-09-13
- AI-powered modular music synthesis written in Rust with full MCP access — simply-chris · 2026-09-13
- Yacine shows third CAD design iteration driven entirely by AI chat from his phone — yacineMTB · 2026-09-13
- CoreWeave Hacks Kicks Off: 200+ Builders Race to Build Self-Correcting Agents in 24 Hours — wandb · 2026-09-13
- Dev rebuilds his 2019 app with Rork, ships it much faster this time — rudrank · 2026-09-13