Independent Tests Show Grok 4.6 (high) Beating 4.7, 23 vs. 19

PawelHuryn · x · 2026-09-22

PawelHuryn ran repeated independent benchmarks comparing Grok 4.6 and 4.7: at high effort 4.6 won 23 vs. 19, and even 4.6 medium scored better. A fourth run at max effort gave 4.7 only a 0.1-point edge — noise. He suspects 4.7 was tuned for popular benchmarks rather than genuinely improved.

Related event: Real-Repo Bug Test: Grok 4.7 Trails GPT-6 and Barely Beats Grok 4.6(4 posts)→

Original post →

More from Models

Models channel →