Bug-Hunting Benchmark: Grok 4.7 Scores 28.7, Trailing GPT-6
Pawel Huryn planted 105 bugs across two real repositories and found GPT-6 Astra (max) leads with 45 points, while Grok 4.7 scored just 28.7, failing to even beat Grok 4.6.
2026-09-22 ~ 2026-09-22 · 2 related posts
- 105 Planted Bugs Put Grok 4.7 at 28.7 vs GPT-6 Astra's 45 in Real-Repo Coding Test — PawelHuryn · 2026-09-22
1 near-duplicate retellings: PawelHuryn