SlopCodeBench Update: GPT-6 Astra Leads at 21%, GLM 5.3 High Hits 19%
teropa · x · 2026-09-06
dexhorthy updates SlopCodeBench with a first multi-model run across thinking-effort settings: GPT-6 Astra Xhigh leads at 21%, GLM 5.3 High reaches 19%, GPT-5.6 Sol XHigh 18%, Fable 5.1 Medium 11%, GPT-5.6 Sol Medium 11%, GLM 5.3 Medium 7%. He notes effort doesn't translate cleanly across providers; the Fable 5.1 Xhigh run waits until next week's quota reset.
More from Models
- "Total OpenAI victory" debate: coder says Cursor beats Codex harness by a wide margin — IndraVahan · 2026-09-06
- GLM Coding Plan ups Flash quotas: unlimited in ZCode, 2x elsewhere — pcuenq · 2026-09-06
- One-line verdict: astra fast ranks clearly above fable 5.1 — NERDDISCO · 2026-09-06
- Researchers dispute Anthropic's NEM reward hacking result as setup artifact — voooooogel · 2026-09-06
- Dev says OpenAI's Astra is first model making progress on his 'unreasonably complex' project — mrjonfinger · 2026-09-06
- GPT-6 Astra turns 38-page cabin blueprints into to-scale 3D walkthrough in ~10 minutes — LukeW · 2026-09-06