HKUST benchmark: even best legal AI agents hallucinate in 89% of trajectories

机器之心 · wechat · 2026-09-26

HKUST researchers released LexAgentHallu, a hierarchical benchmark profiling hallucinations in legal agents across full execution trajectories rather than just final answers. It contains 3,414 instances across 17 legal categories and 6 task types, with a two-tier taxonomy: legal content errors (sources, principles, procedure, fact application) and process errors (planning, memory, tool use), spanning 7 mid-level categories and 27 fine-grained subtypes.

Key findings:

The paper also proposes product defenses: define scope upfront, verify evidence during execution, and check conclusion-reason consistency before output. Limitations: based on Chinese law and single-agent settings. The team is advised by Prof. Han Sihui and Prof. Guo Yike.

Original post →

More from Safety

Safety channel →