Reddit proposes measuring LLMs by cost per accepted task, not cost per token, after Grok 4.7 launch
Crescitaly · reddit · 2026-09-22
xAI announced Grok 4.7 on September 21, claiming better long-running work and self-checking while keeping Grok 4.6's base price and speed.
The author argues that for research and marketing work, leaderboards matter less than a different comparison: give each model the same source-backed brief, then count how much human cleanup is needed—unsupported claims, missed requirements, revision rounds, elapsed time, and total usage—with success criteria fixed in advance. The core point: a cheaper token rate doesn't tell you the cost of a completed task.
The author explicitly notes this is a proposed test, not an experimental result, and doesn't claim Grok wins it.
More from Models
- Asking Claude for a direct quote, it just hands back a hallucinated link — xuenay · 2026-09-22
- Jev vs LLMs: typed probabilistic queries with parallel evaluation instead of token-by-token decoding — techNmak · 2026-09-22
- GPT-6 Astra costs $50 per million output tokens vs DeepSeek's $1.20 — frontier intelligence is cheap if you route tasks right — alex_verem · 2026-09-22
- Xiaomi Releases MiMo-V2.6-Distill-Qwen-9B Distilled Model — Aggravating-Push-207 · 2026-09-22
- Qwen rumored to push 5-10T params with recursive self-improvement, plus Yunqi lineup — lxfater · 2026-09-22
- Xiaomi MiMo 2.6 Pro formalizes Li–Yorke chaos theorem in 6,000+ lines of verified Lean — bookwormengr · 2026-09-22