Hallucination gating reshuffles legal agent benchmark; Grok 4.7 takes the lead

Artificial Analysis and legal AI company Harvey jointly released v1.1 of the LAB-AA legal agent benchmark methodology. The key change is a new hallucination gate: a deliverable only counts toward the headline Hallucination-Gated score if it meets all scoring criteria and contains no "substantive hallucinations." With the gate in place, most model rankings reshuffled dramatically: Grok 4.7 took the top spot, while pre-gate leader Meta Muse Spark 1.3 (max) fell from 26.7% to 8.9%, with over 60% of its passing results found to contain substantive hallucinations. The metric highlights that in legal settings, "getting the task done" and "not making things up" are two independent capabilities.

Confirmed

Why it matters

2026-10-09 ~ 2026-10-09 · 8 related posts

Primary sources