Real telemetry contradicts Grok 4.7 ragebait: 46% fewer tokens per task
ns123abc · x · 2026-09-22
A ragebait YouTuber used leading prompts in a frontend playground to claim Grok 4.7 is 30-80% less token-efficient and 2x more expensive. Actual engineering telemetry says otherwise: Cursor reports only 5% more tokens for median prod requests, and Mercor measured Grok 4.7 using 46% FEWER tokens per task than Grok 4.6 (23.8k vs 44.1k). A case study in vibe-prompt benchmarking vs real-world data.
More from Models
- Limite ships 1B math model trained on under 300B tokens — tensorqt · 2026-09-22
- Best open LLM reportedly cost just $3.5M to post-train — Kyrannio · 2026-09-22
- 37 benchmarks, 130K decisions per model: jev excels at tools and automation — multimodalart · 2026-09-22
- 37 benchmarks, 130K decisions per model: local LLMs tested on one RTX 6000 PRO — multimodalart · 2026-09-22
- Decision Index 0.1: leaderboard asks 130K questions to 30+ open decision models — multimodalart · 2026-09-22
- Xiaomi Releases MiMo-V2.6-Distill-Qwen-9B, a 9B Agentic Model Distilled from Qwen3.5 — anovers · 2026-09-22