HKUST's LexAgentHallu benchmarks legal-agent hallucinations across full execution trajectories, not final answers
jiqizhixin · x · 2026-10-05
Hong Kong University of Science and Technology introduces LexAgentHallu, a benchmark for hallucinations in legal LLM agents.
Key insight: a legal agent can cite a wrong statute and build a fluent argument on top — the more coherent the later reasoning, the easier the initial error hides behind a plausible final answer. Once hallucination enters an execution chain, it becomes the premise for the next legal judgment.
Shift in evaluation: instead of checking only final answers, LexAgentHallu evaluates the full trajectory, localizing errors at specific steps — verifying not just conclusions but whether supporting legal authorities are real, accurate, and consistent across steps.
More from Research
- NeurIPS 2026 accepts Princeton collab paper on learning linear dynamical systems with small memory — HazanPrinceton · 2026-10-05
- Language models map psychedelic drugs' 19 neurotransmitter receptors to subjective experience semantics — danilobzdok · 2026-10-05
- Researcher Predicts a Neural Net Will Prove Deep Learning Works — and No Human Will Understand the Proof — jasondeanlee · 2026-10-05
- Lungfish claims biggest animal genome: 91 billion bases, 30x human DNA — cephaloform · 2026-10-05
- REFRAG beats prompt caching's exact-prefix limit, but 16x compression still won't fit your database in context — CShorten30 · 2026-10-05
- Living Science proposal: papers as living, updatable units of scientific knowledge — ChenhaoTan · 2026-10-05