105 Real Bugs Tested: Grok 4.7 Scores 28.7, No Better Than Grok 4.6; GPT-6 Astra Leads at 45
PawelHuryn · x · 2026-09-22
PawelHuryn planted 105 hard bugs (ones frontier models missed in early 2026) across 2 repos and ran 3 trials per model: GPT-6 Astra (max) led with 45, Muse Spark 1.3 scored 32.2, Grok 4.7 and 4.6 (xhigh) both averaged 28.7, Opus 5 got 27, Qwen3.8-Max 25.7. Grok 4.7 showed no meaningful gain over 4.6 (runs: 25/30/31 vs 27/30/29) — "not the model we've been waiting for." Methodology: unplanted bugs don't count (models over-reporting is arguably negative), cross-family judges score against a secret answer key, incomplete fixes score 0.
Related event: Bug-Hunting Benchmark: Grok 4.7 Scores 28.7, Trailing GPT-6(2 posts)→
More from Models
- Xiaomi open-sources MiMo-V2.6 omni-modal models, topping open-model index at 46.32 — victormustar · 2026-09-22
- SemiAnalysis says open source is dying, yet 20+ open models shipped in the past month — _lewtun · 2026-09-22
- Grok 4.7 posts 59% recall on defensive cyber bench at half the cost of rivals — andreamichi · 2026-09-22
- Grok 4.7 jumps from #9 to #3 on BuildingBench with 0.783, 66% cheaper than Fable 5.1 — ZhitingHu · 2026-09-22
- Jev explained: why the AI community's new favorite isn't a traditional LLM — multiply_matrix · 2026-09-22
- LLMs excel at 1-token output — dev proposes replacing low/medium/high reasoning tiers with token counts — arkuto · 2026-09-22