More tokens, worse scores: Claude models emit up to 562k tokens yet score only 2.8-6.4%

ArtificialAnlys · x · 2026-10-09

In Artificial Analysis's hallucination-gated legal agent benchmark, more output tokens doesn't mean higher scores: GPT-6 Astra (max) scores 8.6% on 81k output tokens per task, under half the 180k of Grok 4.7 (xhigh). The three Claude models generate the most tokens (202k to 562k per task) yet score only 2.8% to 6.4%.

Related event: GPT-6 Astra Tops LAB-AA Legal Benchmark, Output Length Doesn't Help(2 posts)→

Original post →

More from Models

Models channel →