Exploit Bench results are highly harness-sensitive: GLM 5.3 Flash beats 4.1, best in Claude Code
teortaxesTex · x · 2026-10-03
Another GeneralityLabs result shows Exploit Bench performance is very harness-sensitive. V4.1 Flash is far weaker than 5.3 Flash, which performs best in Claude Code with Codex second. The "original" harness is ExploitBench's default (roughly a Python while loop), not ZCode or DSH. The author adds that GLM-5.3 Flash can exceed Mythos Preview's scores, arguing TTC scaling is dissolving categorical capability tiers.
More from Models
- Steve Yegge: Two Weeks With Opus 5.5 — Precision Rivals Fable, Recall Trails on Open-Ended Tasks — Steve_Yegge · 2026-10-03
- OpenAI's Codex global reset appears to skip Business accounts, support suggests buying credits — AdventurousFeeling19 · 2026-10-03
- Mystery 'iguana_necktie' Field Spotted in Anthropic Usage API — bytebot · 2026-10-03
- Insider teases 'new SSI model,' calling Ilya 'a truly remarkable human being' — iruletheworldmo · 2026-10-03
- Reddit user says Opus 5.5 burns weekly cap at 300M tokens, down from 1-2B before — Chemical-Ad-7982 · 2026-10-03
- Meta Muse quality collapsing: inference overload seen behind dropped tasks — brandon_galang · 2026-10-03