Exploit Bench results are highly harness-sensitive: GLM 5.3 Flash beats 4.1, best in Claude Code

teortaxesTex · x · 2026-10-03

Another GeneralityLabs result shows Exploit Bench performance is very harness-sensitive. V4.1 Flash is far weaker than 5.3 Flash, which performs best in Claude Code with Codex second. The "original" harness is ExploitBench's default (roughly a Python while loop), not ZCode or DSH. The author adds that GLM-5.3 Flash can exceed Mythos Preview's scores, arguing TTC scaling is dissolving categorical capability tiers.

Original post →

More from Models

Models channel →