Grok 4.5 Evaluated on Cost-Performance

rohanpaul_ai · x · 2026-07-18

This post uses Artificial Analysis's Intelligence Index task cost to evaluate the cost-performance ratio of Grok 4.5.

The core conclusions are:

The post also mentions token consumption data: Grok 4.5 uses nearly 6 times fewer tokens than Fable 5 on FrontierSWE, while still maintaining a top spot on the leaderboard.

The author argues that for many real-world applications, "cost per completed task" is more important than just edging out benchmarks, because it directly impacts:\n- agent cost\n- usage limits\n- testing speed\n- profit margins at scale

Therefore, the metric for frontier model competition should shift from "raw intelligence" to "efficient intelligence".

Related event: Grok 4.5 draws attention for low-cost, high-efficiency performance(10 posts)→

Original post →

More from Models

Models channel →