SlopCodeBench Update: GPT-6 Astra Leads at 21%, GLM 5.3 High Hits 19%

teropa · x · 2026-09-06

dexhorthy updates SlopCodeBench with a first multi-model run across thinking-effort settings: GPT-6 Astra Xhigh leads at 21%, GLM 5.3 High reaches 19%, GPT-5.6 Sol XHigh 18%, Fable 5.1 Medium 11%, GPT-5.6 Sol Medium 11%, GLM 5.3 Medium 7%. He notes effort doesn't translate cleanly across providers; the Fable 5.1 Xhigh run waits until next week's quota reset.

Original post →

More from Models

Models channel →