Full SlopCodeBench run: GPT-6 Astra edges GPT-5.6 Sol and GLM 5.3, but not by much

cedric_chee · x · 2026-09-13

dexhorthy finished a full SlopCodeBench run (previously only subsets) comparing GPT-6 Astra, GPT-5.6 Sol and GLM 5.3, with Fable 5.1 results pending. Astra scores a few points above GPT-5.5, less than expected, and Sol underperformed. Caveats: ran over a week with provider outages; author suggests aggregating multiple runs for scientific rigor. Context: prior reports of Astra producing gradually bloating code and code-golfing behavior.

Original post →

More from coding & agent

coding & agent channel →