Hallucination gating reshuffles legal agent benchmark; Grok 4.7 takes the lead
Artificial Analysis and legal AI company Harvey jointly released v1.1 of the LAB-AA legal agent benchmark methodology. The key change is a new hallucination gate: a deliverable only counts toward the headline Hallucination-Gated score if it meets all scoring criteria and contains no "substantive hallucinations." With the gate in place, most model rankings reshuffled dramatically: Grok 4.7 took the top spot, while pre-gate leader Meta Muse Spark 1.3 (max) fell from 26.7% to 8.9%, with over 60% of its passing results found to contain substantive hallucinations. The metric highlights that in legal settings, "getting the task done" and "not making things up" are two independent capabilities.
Confirmed
- Ranking shifts: after gating, Grok 4.7 ranks first; Muse Spark 1.3 (max) dropped from first place at 26.7% pre-gate to 8.9%.
- Capability separation: meeting evaluation criteria and avoiding hallucination are distinct skills—Kimi K3 (max) achieved a 93.0% criteria pass rate but averaged 2.09 substantive hallucinations per task.
- Checker selection: six hallucination checkers (including GPT-6 Sol, GPT-6 Luna, and Grok) were benchmarked on deliverables from 8 models over a fixed subset of 20 tasks; GPT-6 Sol (high) was ultimately chosen for production. GPT-6 Sol flagged 470 issues, far more than Grok's 219.
- Cost Pareto frontier: among models with gated pass rate >0, GPT-6 Luna costs $0.22 per task while the Claude family runs $18–22; the frontier also includes GPT-6.1 Sol (max), Muse Spark 1.3 (max), and others.
- Token count doesn't correlate with scores: GPT-6 Astra (max) used 81k output tokens per task for an 8.6% score, below Grok's level; three Claude models output 200k tokens yet still ranked at the bottom.
Why it matters
- Legal deliverables demand extremely high factual accuracy; the hallucination gate elevates "staying error-free" to a core evaluation dimension, giving model selection for professional settings a metric closer to real-world risk.
- The evaluation shows rankings are highly sensitive to how metrics are defined—a reminder not to rely on a single benchmark score, but to look at hallucination rates, cost, output length, and other multi-dimensional signals together.
2026-10-09 ~ 2026-10-09 · 8 related posts
Primary sources
- With hallucination gating, Grok 4.7 tops legal agent benchmark at 9.4% all-pass — ArtificialAnlys ·
- Hallucination gating reshuffles legal benchmark: Meta's Muse Spark falls from 26.7% to 8.9% — ArtificialAnlys ·
- Artificial Analysis and Harvey add hallucination gating to Legal Agent Benchmark (LAB-AA v1.1) — ArtificialAnlys ·
- [source] With hallucination gating, Grok 4.7 tops legal agent benchmark at 9.4% all-pass — ArtificialAnlys · 2026-10-09
- [source] Artificial Analysis and Harvey add hallucination gating to Legal Agent Benchmark (LAB-AA v1.1) — ArtificialAnlys · 2026-10-09
- Six-model hallucination checker bake-off: GPT-6 Sol flags 470 material hallucinations vs Grok's 219 — ArtificialAnlys · 2026-10-09
- [source] Hallucination gating reshuffles legal benchmark: Meta's Muse Spark falls from 26.7% to 8.9% — ArtificialAnlys · 2026-10-09
- Passing criteria vs. not hallucinating: Kimi K3 hits 93% pass rate with 2.09 hallucinations per task — ArtificialAnlys · 2026-10-09
- Legal benchmark Pareto frontier: GPT-6 Luna at $0.22/task vs Claude at $18-22/task — ArtificialAnlys · 2026-10-09
- More tokens, worse scores: Claude models emit up to 562k tokens yet score only 2.8-6.4% — ArtificialAnlys · 2026-10-09
- Harvey LAB-AA eval: GPT-6 Astra tops scores while Grok 4.7 outputs 2x its tokens — ArtificialAnlys · 2026-10-09