Grok 4.7 Slammed in Benchmarks: 2.5x Tokens for Just 2 Points More
ivan_bezdomny · x · 2026-09-22
xAI pitched Grok 4.7 as a notable improvement over 4.6 at the same price and speed, but third-party comparisons tell a different story:
- Grok 4.7 xhigh burns 2.5x the tokens of 4.6 high to score just 2 points higher on artificial arena
- On cursorbench, 4.7 low scores below Grok 4.6 low while using more tokens and more steps
The poster concludes the model is "dead on arrival," questioning whether the official claims hold up in practice.
More from Models
- Xiaomi's MiMo-V2.6-Flash-RL trends on Hugging Face with multimodal and agent skills — XiaomiMiMo · 2026-09-22
- Blogger corrects himself: the real surprise is MiMo-2.6, beating Grok 4.7 at much lower cost — kimmonismus · 2026-09-22
- cHHillee defends Tinker: you can own your post-training codebase, only trainer and sampler are abstracted — cHHillee · 2026-09-22
- Unverified DataBench Charts Fuel Rumors of OpenAI's Internal Model 'Luna' Ahead of GPT-6 — almmaasoglu · 2026-09-22
- Mimo 2.6 Pro hands-on: needs more steering, rivals top models after corrections — power97992 · 2026-09-22
- Speculation: Grok Pro line is an extension of Flash line, mxfp4 QAT likely speeds RL rollouts — stochasticchasm · 2026-09-22