Reddit proposes measuring LLMs by cost per accepted task, not cost per token, after Grok 4.7 launch

Crescitaly · reddit · 2026-09-22

xAI announced Grok 4.7 on September 21, claiming better long-running work and self-checking while keeping Grok 4.6's base price and speed.

The author argues that for research and marketing work, leaderboards matter less than a different comparison: give each model the same source-backed brief, then count how much human cleanup is needed—unsupported claims, missed requirements, revision rounds, elapsed time, and total usage—with success criteria fixed in advance. The core point: a cheaper token rate doesn't tell you the cost of a completed task.

The author explicitly notes this is a proposed test, not an experimental result, and doesn't claim Grok wins it.

Original post →

More from Models

Models channel →