Grok 4.7 tops Harvey LAB-AA v1.1, beating Claude Opus 5.5 and GPT-6 Astra

XFreeze · x · 2026-10-09

XFreeze reports that Grok 4.7 ranked #1 on the Harvey LAB-AA v1.1 leaderboard, outperforming Claude Opus 5.5, GPT-6 Astra, Muse Spark 1.3 and others. The benchmark is unusually harsh: a single material hallucination zeroes the entire task, so it measures end-to-end task completion without substantive hallucinations rather than just answer accuracy — evidence, the author argues, that Grok 4.7 is excelling at high-stakes reasoning.

Related event: Grok 4.7 tops Harvey LAB-AA after hallucination gating reshuffles legal agent rankings(9 posts)→

Original post →

More from Models

Models channel →