With hallucination gating, Grok 4.7 tops legal agent benchmark at 9.4% all-pass

ArtificialAnlys · x · 2026-10-09

Artificial Analysis and Harvey released LAB-AA v1.1, adding a hallucination check to the Legal Agent Benchmark: a task only counts toward the headline Hallucination-Gated All-Pass Rate if deliverables meet every rubric criterion with no material hallucinations.

Key findings:

Related event: Hallucination gating reshuffles legal agent benchmark; Grok 4.7 takes the lead(8 posts)→

Original post →

More from Models

Models channel →