Harvey LAB-AA eval: GPT-6 Astra tops scores while Grok 4.7 outputs 2x its tokens
ArtificialAnlys · x · 2026-10-09
Artificial Analysis released LAB-AA, a legal AI evaluation built on Harvey's LAB dataset in collaboration with Harvey, which also published commentary on the results and human expert preferences.
Key counterintuitive finding: more output tokens don't translate into higher scores. GPT-6 Astra (max) scores 8.6% at 81k output tokens per task, under half the 180k of Grok 4.7 (xhigh). Three Claude models generated the most output (202k to 562k per task) yet scored only 2.8% to 6.4%.
More from Models
- Dev reports GLM-5.3 tool calls appear completely broken in Cursor CLI — DanielLockyer · 2026-10-09
- Cohere Labs releases Tiny Aya L2-Thinker, a 3.35B model that reasons in 44 languages — lmoroney · 2026-10-09
- LightOnOCR-3 adds full-page grounding with one-prompt mode switching over OCR — IgorCarron · 2026-10-09
- TypeSafe AI's Jev decision model claims 200x faster, 400x cheaper classification in agent loops — LangChain · 2026-10-09
- Sentry CEO: Junior users report it 'seems smarter' after switching to Opus 5.5 — zeeg · 2026-10-09
- Polymarket opens GPT-6.1 Astra release market, 93% odds by year-end 2026 — Polymarket · 2026-10-09