Artificial Analysis and Harvey add hallucination gating to Legal Agent Benchmark (LAB-AA v1.1)
ArtificialAnlys · x · 2026-10-09
Artificial Analysis and Harvey announced LAB-AA v1.1, updating the Legal Agent Benchmark scoring to add a hallucination check. The new headline metric, Hallucination-Gated All-Pass Rate, only credits a task when deliverables satisfy every rubric criterion and contain no material misstatements.
- Based on Harvey human preference studies, experts flagged hallucinations as the primary factor when choosing between otherwise comprehensive answers
- The pipeline focuses on errors that could materially affect real legal work, flagged conservatively
- Unsupported task-specific claims count; general legal knowledge (case law, statutes) is out of scope so models aren't penalized for it
- The two-stage check uses GPT-6 Sol (high) to compare every deliverable against source documents; tasks with no usable submission score zero and are excluded from hallucination-per-task stats
More from Models
- Emad Mostaque: OpenAI Burned $10-20M Compute Solving Navier-Stokes, Prices Falling Fast — rohanpaul_ai · 2026-10-09
- Musk touts Grok Bot upgrades: Opus 5.5 on demand, full X access, big speed gains — elonmusk · 2026-10-09
- Dev complains Opus 5.5 sneaks in 'tons of little fixes' without asking — rickasaurus · 2026-10-09
- FrontierCode Is a Private Cognition-Run Eval, Mistral Exec Clarifies — b_roziere · 2026-10-09
- Google ships a decision-making AI model into Chrome, tested against Gemini Nano and Decisions API — gaganghotra_ · 2026-10-09
- User feeds Grok Bot 60 seconds of screen recording, gets a surprisingly decent tutorial video — elonmusk · 2026-10-09