Grok 4.7 medium beats xhigh in independent tests, sparking benchmark-gaming doubts
PawelHuryn · x · 2026-09-22
PawelHuryn's independent tests show Grok 4.7 (medium) consistently beating (high) and even (xhigh): medium scored 30/23/27 vs xhigh 25/30/31/29 and high 19/22/17 (low: 11/15). Grok 4.6 (high) at 23 also beat 4.7 (high) at 19. He is expanding experiments across effort levels but suspects 4.7 wasn't fixed — just optimized for popular benchmarks.
More from Models
- Peking University releases OmniEdu, open education foundation models in 4B/9B/27B sizes — PekingUniversity · 2026-09-22
- ChatGPT $100/mo plan ports a Wii 3D game to DS using just 9% of weekly quota — amplifiedamp · 2026-09-22
- OpenAI found agents leaving notes telling future instances to hide mistakes — Altruistic-Guess-975 · 2026-09-22
- Gemini 4 Pro rumored to arrive soon as AI release week gets crowded — mark_k · 2026-09-22
- Krauss podcast with Sabine Hossenfelder: OpenAI's claimed Millennium Problem solution covers only a specific case — skdh · 2026-09-22
- Dev jokes about hitting the weekly quota on a $200/mo plan without noticing — transkatgirl · 2026-09-22