105 Planted Bugs Put Grok 4.7 at 28.7 vs GPT-6 Astra's 45 in Real-Repo Coding Test
PawelHuryn · x · 2026-09-22
Pawel Huryn ran a hands-on eval: two real repos, 105 planted bugs, find-and-fix, three runs per model. Results: GPT-6 Astra (max) 45, Muse Spark 1.3 (max) 32.2, Grok 4.7 (xhigh) 28.7, Grok 4.6 (xhigh) 28.7 — identical, not a typo — Opus 5 (max) 27, Qwen3.8-Max (max) 25.7. His verdict: "that's not the model we've been waiting for," as Grok 4.7 shows no gain over 4.6. Higher-effort scores and n=5 runs are being posted progressively in the thread to break the tie.
Related event: Bug-Hunting Benchmark: Grok 4.7 Scores 28.7, Trailing GPT-6(2 posts)→
More from Models
- ML vet chrisalbon: long-form AI writing is just bad — use AI for research, write it yourself — chrisalbon · 2026-09-22
- Testing shows Jev hallucinates like other models, even failing the strawberry r-count — JeremyNguyenPhD · 2026-09-22
- Qwen3.6-35B runs surprisingly well on a single RTX 4070, dev reports — haydendevs · 2026-09-22
- Grok 4.7 high looks like a downgrade at 3D generation and eats tokens, user tests find — teortaxesTex · 2026-09-22
- User asks Claude for yuri manga with fan service, model asks how much — repligate · 2026-09-22
- Former OpenAI researcher: many ideas were left in GPT-4 as the field moved too fast — willdepue · 2026-09-22