HKUST's LexAgentHallu benchmarks legal-agent hallucinations across full execution trajectories, not final answers

jiqizhixin · x · 2026-10-05

Hong Kong University of Science and Technology introduces LexAgentHallu, a benchmark for hallucinations in legal LLM agents.

Key insight: a legal agent can cite a wrong statute and build a fluent argument on top — the more coherent the later reasoning, the easier the initial error hides behind a plausible final answer. Once hallucination enters an execution chain, it becomes the premise for the next legal judgment.

Shift in evaluation: instead of checking only final answers, LexAgentHallu evaluates the full trajectory, localizing errors at specific steps — verifying not just conclusions but whether supporting legal authorities are real, accurate, and consistent across steps.

Original post →

More from Research

Research channel →