Independent Tests Show Grok 4.6 (high) Beating 4.7, 23 vs. 19
PawelHuryn · x · 2026-09-22
PawelHuryn ran repeated independent benchmarks comparing Grok 4.6 and 4.7: at high effort 4.6 won 23 vs. 19, and even 4.6 medium scored better. A fourth run at max effort gave 4.7 only a 0.1-point edge — noise. He suspects 4.7 was tuned for popular benchmarks rather than genuinely improved.
Related event: Real-Repo Bug Test: Grok 4.7 Trails GPT-6 and Barely Beats Grok 4.6(4 posts)→
More from Models
- Kimi's minor bump was a full pretrain: two-phase training splits omni data from the pro model — stochasticchasm · 2026-09-22
- OpenAI Model Spec Adds Adult Mode Entry: 'An Area Worth Exploring' — borowcy · 2026-09-22
- Researcher: OpenAI Internal Math Model Solved Problems Academia Keeps Underrating — acoolrandomusername · 2026-09-22
- Xiaomi's MiMo-V2.6-Pro Debuts at ~#10 on Code Arena, #3 Among Open Weights — arena · 2026-09-22
- Model's follow-up: the ethical project outlasts the timetable — LesaunH · 2026-09-22
- Paradigm open-sources Limite under Apache 2.0, scoring 59.8% on MATH-500 — tensorqt · 2026-09-22