105 real bugs: GLM-5.3 fixes 19 vs Grok 4.7's 27, at half the speed
PawelHuryn · x · 2026-08-26
Testing GLM-5.3 on the author's bug hunt bench — 2 repos, 105 real bugs — it solved 19 vs 27 for Grok 4.7, while running 2x slower. Notably, 49 of 105 bugs were never fixed by any of the 16 frontier models tested; when that happens, "bugs are solved." The author also notes Grok 4.6 is not on the cost Pareto frontier: Luna (max) is an incredibly powerful bug hunter for its price. Live benchmark and evidence at bughunt.productcompass.pm.
Related event: GLM-5.3 Trails Grok 4.7 in 105-Real-Bug Benchmark(2 posts)→
More from Models
- View: Tokens-per-second matters more than model size now — natesiggard · 2026-08-26
- Tiel-Coder-35B achieves 121.4 tok/s for local inference — DerTomsn · 2026-08-26
- Benchmark: Tool Calling Performance of Qwen 35B-A3B Variants — OsmanthusBloom · 2026-08-26
- Grok 4.6 now available on OpenCode Go subscription — veggie_eric · 2026-08-26
- One Claude Design Task Burned 95% of the $100/Mo Plan — zeeg · 2026-08-26
- Claude Team Premium burns 3x faster than Max 5x plans — james_mtc · 2026-08-26