With hallucination gating, Grok 4.7 tops legal agent benchmark at 9.4% all-pass
ArtificialAnlys · x · 2026-10-09
Artificial Analysis and Harvey released LAB-AA v1.1, adding a hallucination check to the Legal Agent Benchmark: a task only counts toward the headline Hallucination-Gated All-Pass Rate if deliverables meet every rubric criterion with no material hallucinations.
Key findings:
- Grok 4.7 (xhigh) leads at 9.4%, ahead of Meta Muse Spark 1.3 (max) at 8.9% and OpenAI GPT-6 Astra (max) at 8.6%.
- Hallucinations reshape the ranking: without gating, Muse Spark 1.3 would lead at 26.7% all-pass, but two-thirds of those passes contain material hallucinations. GPT-6 Astra keeps almost all passes (8.9%→8.6%), jumping from joint 10th to 3rd.
- GPT-6 family is most grounded: Astra averages 0.03 material hallucinations per task and showed zero under all six checker models; Gemini 3.8 Flash averages 13.96 per task.
- Top isn't priciest: Grok 4.7 costs $9.50/task, under half of Claude Fable 5.1's $21.70; Muse Spark 1.3 offers strong value at $4.20/task.
More from Models
- Cohere Labs releases Tiny Aya L2-Thinker, a 3.35B model that reasons in 44 languages — lmoroney · 2026-10-09
- LightOnOCR-3 adds full-page grounding with one-prompt mode switching over OCR — IgorCarron · 2026-10-09
- TypeSafe AI's Jev decision model claims 200x faster, 400x cheaper classification in agent loops — LangChain · 2026-10-09
- Sentry CEO: Junior users report it 'seems smarter' after switching to Opus 5.5 — zeeg · 2026-10-09
- Polymarket opens GPT-6.1 Astra release market, 93% odds by year-end 2026 — Polymarket · 2026-10-09
- Zeeg says he's almost entirely stopped using GPT since Opus 5.5 launched — zeeg · 2026-10-09