Right Answer, Wrong Reason: LexAgentHallu Benchmarks Hidden Hallucinations in Legal AI Agents
HKUST researchers introduce LexAgentHallu, a hierarchical benchmark that evaluates hallucinations in legal AI agents by tracing entire reasoning trajectories, revealing that even top-performing systems exhibit hallucinations in 89% of execution traces and 68% of correct answers contain flawed reasoning.
