105 Real Bugs Tested: Grok 4.7 Scores 28.7, No Better Than Grok 4.6; GPT-6 Astra Leads at 45

PawelHuryn · x · 2026-09-22

PawelHuryn planted 105 hard bugs (ones frontier models missed in early 2026) across 2 repos and ran 3 trials per model: GPT-6 Astra (max) led with 45, Muse Spark 1.3 scored 32.2, Grok 4.7 and 4.6 (xhigh) both averaged 28.7, Opus 5 got 27, Qwen3.8-Max 25.7. Grok 4.7 showed no meaningful gain over 4.6 (runs: 25/30/31 vs 27/30/29) — "not the model we've been waiting for." Methodology: unplanted bugs don't count (models over-reporting is arguably negative), cross-family judges score against a secret answer key, incomplete fixes score 0.

Related event: Bug-Hunting Benchmark: Grok 4.7 Scores 28.7, Trailing GPT-6(2 posts)→

Original post →

More from Models

Models channel →