Right Answer, Wrong Reason: LexAgentHallu Benchmarks Hidden Hallucinations in Legal AI Agents
HKUST researchers introduce LexAgentHallu, a hierarchical benchmark that evaluates hallucinations in legal AI agents by tracing entire reasoning trajectories, revealing that even top-performing systems exhibit hallucinations in 89% of execution traces and 68% of correct answers contain flawed reasoning.
Hong Kong University of Science and Technology (HKUST) researchers have introduced LexAgentHallu , a hierarchical benchmark designed to profile hallucinations in legal AI agents by examining the complete execution trajectory rather than only the final answer. The benchmark addresses a critical gap: as legal large language models evolve into agents that plan, retrieve, read, and act across multiple steps, errors in early steps — such as citing an incorrect legal source, miscalculating a procedural deadline, or misinterpreting a tool's output — can propagate silently through the reasoning chain and be masked by a seemingly correct final answer.
Benchmark Design: Two-Layer Taxonomy and Verifiable Construction
LexAgentHallu organizes hallucinations into a two-layer taxonomy :
Layer 1 (Legal Content) : errors in legal sources, legal principles, procedural rules, and fact-rule mapping.
Layer 2 (Agent Process) : errors in planning & reasoning, memory, and tool invocation & observation understanding.
Together the two layers comprise 7 mid-level categories and 27 fine-grained subcategories . The taxonomy is intended not merely to rename errors but to enable precise localization — for example, source fabrication calls for citation verification, fact-rule mismatch requires element re-checking, tool parameter errors demand interface constraints, and memory drift points to context management.
The benchmark was built through a four-stage pipeline (Figure 1 in the paper): data collection from diverse legal tasks, filtering of easy samples, annotation by legally trained annotators who verified answers, sources, and reasoning traces, and expert review. The final dataset contains 3,414 instances covering 17 legal categories and 6 task types (legal reasoning, legal consultation, legal knowledge QA, judgment analysis, case analysis, judgment prediction).
Evaluation Results: Pervasive Hallucinations Across Configurations
The paper evaluates 18 closed-source and open-source agent configurations . Key findings:
Even the best-performing configuration still exhibits hallucinations in 89% of its execution traces.
All systems show legal-content hallucination frequencies no lower than 87.5% .
Agent architecture influences error distribution: ReAct-style action-observation loops perform better on overall and legal-content hallucinations; legal-specific workflows reduce process errors more effectively; plan-then-execute does not show consistent superiority in this experiment.
The results underscore that model, tool, and orchestration must be evaluated as an integrated system — a stronger base model can amplify errors in an unsuitable retrieval pipeline, while a smaller model with stricter tool protocols and process constraints may achieve more stable behavior.
RAWR Metric: Right Answer, Wrong Reason
The authors propose the RAWR (Right-Answer-Wrong-Reason) metric, which isolates samples where the final answer is correct and then checks whether the execution trace still contains hallucinations. Results show:
Under the correct-answer condition, 68% of traces contain at least one legal-content hallucination .
37% of traces contain at least one agent-process hallucination .
This demonstrates that a correct final answer does not guarantee a clean, verifiable reasoning chain — a crucial distinction for legal counseling, contract review, and litigation support where users need traceable, auditable justifications.
Task-Type Variation and Error Co-occurrence Patterns
Hallucination frequencies differ markedly across task types (Figure 3):
Judgment analysis reaches 1.00 legal-content hallucination frequency; case analysis 0.95.
Process hallucinations are highest in judgment analysis (0.83) and judgment prediction (0.71) .
Open-ended, long-horizon tasks requiring multi-step evidence integration, deadline judgment, or rule application are especially prone to both content and process errors.
Analysis of the 27×27 subcategory co-occurrence matrix (Figure 4) reveals structured error combinations rather than random scatter. Notable lift scores:
Procedural deadline errors co-occur with remedy-path errors at lift 2.82 .
Procedural step errors co-occur with procedural consequence errors at lift 2.59 .
Cross-group signal: legal-source hierarchy errors co-occur with legal-stance confusion at lift 7.07 .
These patterns indicate that many seemingly separate errors share a common upstream cause — e.g., a hierarchy misjudgment cascades into stance and fact-application errors; a deadline miscalculation propagates through steps and remedies. For system governance, high-risk errors (source hierarchy, deadline calculation, procedural steps, fact-element mapping) can serve as early-warning signals to trigger re-retrieval, rule verification, or human review before the error chain advances.
From Benchmark to Product Guardrails
The paper translates the benchmark insights into three concrete guardrails for deployed legal agents:
Pre-generation: clarify boundaries — confirm applicable jurisdiction, task scope, and answerability; refuse definitive judgments when jurisdiction, statute of limitations, or procedural path are unclear.
During execution: verify evidence — continuously cross-check retrieval results, citation sources, key facts, and tool return values; an intermediate result must not be passed forward merely because it “looks reasonable.”
Pre-output: check consistency — ensure conclusions and reasons mutually support each other, specifically validating source authority, procedural deadlines, constituent elements, and remedy paths; a correct final answer does not replace this step.
More broadly, the evaluation unit should expand from a single model to the model–tool–workflow triad : how the system retrieves, when it reflects, how it retains intermediate state, and under what conditions it stops are all part of reliability.
Limitations and Scope
The authors caution that LexAgentHallu is primarily grounded in the Chinese legal system and focuses on single-agent scenarios . Transfer to common-law jurisdictions, multi-jurisdiction retrieval, or multi-agent collaboration would require redefining source authority, precedent applicability, and procedural rules.
Reference : Zhou, Y., Zheng, M., Cao, C., et al. (2024). LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents . arXiv:2609.09754. https://arxiv.org/abs/2609.09754
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
