Nat Lambert flags Opus as top token-gobbler: leading the Intelligence Index is a red flag

natolambert · x · 2026-09-23

A quoted post notes the new Opus consumes the most tokens per Intelligence Index task of any model benchmarked by Artificial Analysis, which explains Anthropic's price cut. RL researcher Nathan Lambert adds that models topping such leaderboards are usually a red flag, hinting high scores may come from brute-force token usage rather than genuine efficiency. He'd love to be proven wrong.

Related event: Claude Opus 5.5 Launches to Top Rankings with Lower Price and Faster Speed(28 posts)→

Original post →

More from Models

Models channel →