Full SlopCodeBench run: GPT-6 Astra edges GPT-5.6 Sol and GLM 5.3, but not by much
cedric_chee · x · 2026-09-13
dexhorthy finished a full SlopCodeBench run (previously only subsets) comparing GPT-6 Astra, GPT-5.6 Sol and GLM 5.3, with Fable 5.1 results pending. Astra scores a few points above GPT-5.5, less than expected, and Sol underperformed. Caveats: ran over a week with provider outages; author suggests aggregating multiple runs for scientific rigor. Context: prior reports of Astra producing gradually bloating code and code-golfing behavior.
More from coding & agent
- Yacine shows third CAD design iteration driven entirely by AI chat from his phone — yacineMTB · 2026-09-13
- CoreWeave Hacks Kicks Off: 200+ Builders Race to Build Self-Correcting Agents in 24 Hours — wandb · 2026-09-13
- Dev rebuilds his 2019 app with Rork, ships it much faster this time — rudrank · 2026-09-13
- Collage app screenshots into a contact sheet to slash your AI agent's context usage — pvncher · 2026-09-13
- Ruff author Charlie Marsh: we overestimate human code quality — especially our own — charliermarsh · 2026-09-13
- Sourcegraph's Amp goes free with BYOK, dropping all limits and fees — HamelHusain · 2026-09-13