More tokens, worse scores: Claude models emit up to 562k tokens yet score only 2.8-6.4%
ArtificialAnlys · x · 2026-10-09
In Artificial Analysis's hallucination-gated legal agent benchmark, more output tokens doesn't mean higher scores: GPT-6 Astra (max) scores 8.6% on 81k output tokens per task, under half the 180k of Grok 4.7 (xhigh). The three Claude models generate the most tokens (202k to 562k per task) yet score only 2.8% to 6.4%.
Related event: GPT-6 Astra Tops LAB-AA Legal Benchmark, Output Length Doesn't Help(2 posts)→
More from Models
- Musk Endorses Grok-Powered X Bot After Users Call It Surprisingly Good — elonmusk · 2026-10-09
- Blogger predicts Fable and Astra are 2+ weeks away, giving Gemini 4 Argon a window — bindureddy · 2026-10-09
- Try LightOnOCR-3 on your hardest document: early user demos circulate — IgorCarron · 2026-10-09
- When OpenAI and Claude refuse to reverse-engineer Sonos, this dev turns to Kimi — doodlestein · 2026-10-09
- Leak: GPT-6.1 Sol 'Ultrafast' Rolls Out at $12/M Input, 8x Standard Speed — testingcatalog · 2026-10-09
- Legal benchmark Pareto frontier: GPT-6 Luna at $0.22/task vs Claude at $18-22/task — ArtificialAnlys · 2026-10-09