HaluMem: Evaluating Hallucinations in AI Agent Memory Systems
HaluMem introduces an operation-level benchmark for AI agent memory systems, evaluating extraction, update, and QA stages to uncover hallucinations; experiments on six systems show high recall often accompanies false memories, while low update hallucination rates mask massive omissions, and upstream memory errors propagate to final answers.
Background: The Hidden Problem Behind Wrong Answers
When an AI assistant misstates a user's new job title, the error is often attributed to the model "hallucinating" at answer time. However, the mistake may originate earlier: during memory extraction, updating, or retrieval. The same wrong answer can stem from fundamentally different memory failures.
HaluMem: An Operation-Level Hallucination Benchmark
Researchers from China Telecom Research Institute and MemTensor propose HaluMem , the first benchmark evaluating hallucinations in agent memory systems at the operation level. Accepted at EMNLP 2026, HaluMem checks three stages with ground-truth labels:
Memory Extraction : Are relevant facts recorded completely and accurately? Are fabricated details inserted?
Memory Update : When new information arrives, are old memories correctly modified? Are updates missed, wrong, or conflicting?
Memory QA : Can the system answer correctly using memory? How many errors are due to missing information vs. wrong retrieval?
Unlike end-to-end QA benchmarks, HaluMem triggers checks immediately after each dialogue session, providing per-stage diagnostics.
Data Construction: Simulated Lifespans with Million-Token Contexts
HaluMem uses a six-stage, user-centric pipeline: user profile → life skeleton → event stream → session summaries & memory points → multi-turn dialogues → evaluation questions. This simulates 10–20 year user lifespans with evolving careers, health, interests, and relationships, forming three memory types: profile, event, relation. Updates preserve old/new correspondence for verifiable ground truth.
Dialogues include assistant hallucinations (unconfirmed details), implicit expressions, and coreference. Two versions:
HaluMem-Medium : 20 users, 30,073 dialogue turns, ~160k tokens/user, 14,948 memory points, 3,467 QA pairs.
HaluMem-Long : Same core memories, but irrelevant dialogue inserted within and between sessions, expanding context to ~1M tokens/user, 53,516 total turns. The Long version adds distraction, not new facts, enabling study of retention under noise.
QA covers six types: basic recall, multi-hop reasoning, dynamic update, memory boundary, generalization, conflict detection. Human evaluation of 700 Medium sessions achieved 95.70% correctness.
Experiments: Six Memory Systems Tested
Evaluated systems: Mem0, Mem0-Graph, Memobase, MemOS, Supermemory, Zep. Unified settings: top-10 memories for update verification, top-20 for QA, GPT-4o for answer generation. Zep omitted from extraction metrics due to API limits.
Finding 1: More Memories ≠ More Reliable Memories
On Medium, extraction recall: Mem0 42.91%, Mem0-Graph 43.28%, Memobase 14.55%. On Long, recall drops drastically: 3.23%, 2.24%, 6.18% — core memories unchanged, only noise added.
Conversely, MemOS (81.90% recall on Long) and Supermemory (53.02%) cover more targets but suffer low False Memory Resistance (FMR): 28.85% and 36.86% respectively. FMR measures ability to ignore assistant-invented, unconfirmed details. High recall does not automatically confer strong false-memory filtering.
Reliable memory requires both: don't miss what should be recorded, don't store what shouldn't be trusted.
Finding 2: Low Update Hallucination Rate Masks Massive Omissions
Update hallucination rate (errors among attempted updates) is <1.2% for all systems on both versions. However, update omission rates on Long exceed 60% for five of six systems. Mem0 and Mem0-Graph omission rates: 98.51% and 98.40%, yielding correct update rates of only 1.45% and 1.47%. MemOS leads with 65.25% correct updates but still misses ~1/3.
Why low hallucination coexists with high omission? Because hallucination rate is computed only over attempted updates; if few updates are attempted, the error proportion stays low. Moreover, updates depend on prior extraction: if a fact was never stored, it cannot be updated. Operation-level evaluation reveals that update failures often originate at the first write.
Finding 3: Upstream Memory Errors Propagate to Final Answers
Under unified QA, no system reaches 70% accuracy. MemOS tops at 67.23% (Medium) and 64.44% (Long). Mem0 drops from 53.02% to 28.11% accuracy, while its omission rate rises from 27.81% to 54.60%, mirroring its extraction recall collapse.
Question-type analysis: systems handle unknown-info and false-premise detection relatively well, but struggle with multi-hop reasoning, dynamic updates, and generalization. Final answers reflect the entire memory pipeline; improving them requires fixing extraction, update, and retrieval jointly.
Robustness check: re-evaluating 10 users with GPT-5.4-mini as judge changes absolute scores but preserves rankings, trade-offs, and core findings.
From "Can Remember" to "Remembers Reliably"
HaluMem makes improvement directions concrete: if extraction fails, check coverage and source credibility; if update fails, verify old-new linking and completion; if QA fails, trace back through retrieved memories and upstream records. Single metrics are misleading: high recall may admit false memories, low update hallucination may hide omissions, and final QA scores can mask internal failures.
Limitations: current data focuses on profile/event/relation memories, not tool/skill memories; system API openness affects observability. Nevertheless, HaluMem provides a reusable methodology: verify each memory operation behind the final answer, turning "where did it go wrong?" into a measurable question.
Returning to the opening example: next time the assistant gets the job title wrong, we can ask: was it recorded correctly initially? Was it updated after promotion? Which record was used at answer time? When these questions are answerable, AI moves closer to truly reliable long-term memory.
Paper: https://arxiv.org/pdf/2511.03506 Open-source:
https://github.com/MemTensor/HaluMemSigned-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
