105 real bugs: GLM-5.3 fixes 19 vs Grok 4.7's 27, at half the speed

PawelHuryn · x · 2026-08-26

Testing GLM-5.3 on the author's bug hunt bench — 2 repos, 105 real bugs — it solved 19 vs 27 for Grok 4.7, while running 2x slower. Notably, 49 of 105 bugs were never fixed by any of the 16 frontier models tested; when that happens, "bugs are solved." The author also notes Grok 4.6 is not on the cost Pareto frontier: Luna (max) is an incredibly powerful bug hunter for its price. Live benchmark and evidence at bughunt.productcompass.pm.

Related event: GLM-5.3 Trails Grok 4.7 in 105-Real-Bug Benchmark(2 posts)→

Original post →

More from Models

Models channel →