Nat Lambert flags Opus as top token-gobbler: leading the Intelligence Index is a red flag
natolambert · x · 2026-09-23
A quoted post notes the new Opus consumes the most tokens per Intelligence Index task of any model benchmarked by Artificial Analysis, which explains Anthropic's price cut. RL researcher Nathan Lambert adds that models topping such leaderboards are usually a red flag, hinting high scores may come from brute-force token usage rather than genuine efficiency. He'd love to be proven wrong.
Related event: Claude Opus 5.5 Launches to Top Rankings with Lower Price and Faster Speed(28 posts)→
More from Models
- ProgramBench: rebuilding programs from binaries is brutal — Claude Opus 5 leads at 4.5% resolved — jyangballin · 2026-09-23
- GPT-6 Sol and Luna hit Arena, plus a matched head-to-head vs GPT-5.6 Sol — arena · 2026-09-23
- Plinius leaks full Claude Opus-5.5 system prompt, over 1.9M characters with tools — ivan_bezdomny · 2026-09-23
- Claude Opus 5.5 Frontend Tests: Suspected Quantized fable 5.1, Stable but Heavier Reasoning — karminski3 · 2026-09-23
- Opus 5.5, GPT-6 Sol and Luna drop the same night as AI pace debate rages — ThePeterMick · 2026-09-23
- Travel planner PlanMyVisit switches to GPT-6: faster and half the cost — alexbainbridge · 2026-09-23