105 Planted Bugs Put Grok 4.7 at 28.7 vs GPT-6 Astra's 45 in Real-Repo Coding Test

PawelHuryn · x · 2026-09-22

Pawel Huryn ran a hands-on eval: two real repos, 105 planted bugs, find-and-fix, three runs per model. Results: GPT-6 Astra (max) 45, Muse Spark 1.3 (max) 32.2, Grok 4.7 (xhigh) 28.7, Grok 4.6 (xhigh) 28.7 — identical, not a typo — Opus 5 (max) 27, Qwen3.8-Max (max) 25.7. His verdict: "that's not the model we've been waiting for," as Grok 4.7 shows no gain over 4.6. Higher-effort scores and n=5 runs are being posted progressively in the thread to break the tie.

Related event: Bug-Hunting Benchmark: Grok 4.7 Scores 28.7, Trailing GPT-6(2 posts)→

Original post →

More from Models

Models channel →