Harvey LAB-AA eval: GPT-6 Astra tops scores while Grok 4.7 outputs 2x its tokens

ArtificialAnlys · x · 2026-10-09

Artificial Analysis released LAB-AA, a legal AI evaluation built on Harvey's LAB dataset in collaboration with Harvey, which also published commentary on the results and human expert preferences.

Key counterintuitive finding: more output tokens don't translate into higher scores. GPT-6 Astra (max) scores 8.6% at 81k output tokens per task, under half the 180k of Grok 4.7 (xhigh). Three Claude models generated the most output (202k to 562k per task) yet scored only 2.8% to 6.4%.

Related event: Hallucination gating reshuffles legal agent benchmark; Grok 4.7 takes the lead(8 posts)→

Original post →

More from Models

Models channel →