HaluMem: Evaluating Hallucinations in AI Agent Memory Systems

HaluMem introduces an operation-level benchmark for AI agent memory systems, evaluating extraction, update, and QA stages to uncover hallucinations; experiments on six systems show high recall often accompanies false memories, while low update hallucination rates mask massive omissions, and upstream memory errors propagate to final answers.

Machine Heart
Machine Heart
Machine Heart
HaluMem: Evaluating Hallucinations in AI Agent Memory Systems

Background: The Hidden Problem Behind Wrong Answers

When an AI assistant misstates a user's new job title, the error is often attributed to the model "hallucinating" at answer time. However, the mistake may originate earlier: during memory extraction, updating, or retrieval. The same wrong answer can stem from fundamentally different memory failures.

HaluMem: An Operation-Level Hallucination Benchmark

Researchers from China Telecom Research Institute and MemTensor propose HaluMem , the first benchmark evaluating hallucinations in agent memory systems at the operation level. Accepted at EMNLP 2026, HaluMem checks three stages with ground-truth labels:

Memory Extraction : Are relevant facts recorded completely and accurately? Are fabricated details inserted?

Memory Update : When new information arrives, are old memories correctly modified? Are updates missed, wrong, or conflicting?

Memory QA : Can the system answer correctly using memory? How many errors are due to missing information vs. wrong retrieval?

Unlike end-to-end QA benchmarks, HaluMem triggers checks immediately after each dialogue session, providing per-stage diagnostics.

Figure 1: Comparison of HaluMem with existing evaluation methods. Besides checking final answer, also checks if memory deviates during extraction and update. (Paper Figure 2)
Figure 1: Comparison of HaluMem with existing evaluation methods. Besides checking final answer, also checks if memory deviates during extraction and update. (Paper Figure 2)

Data Construction: Simulated Lifespans with Million-Token Contexts

HaluMem uses a six-stage, user-centric pipeline: user profile → life skeleton → event stream → session summaries & memory points → multi-turn dialogues → evaluation questions. This simulates 10–20 year user lifespans with evolving careers, health, interests, and relationships, forming three memory types: profile, event, relation. Updates preserve old/new correspondence for verifiable ground truth.

Dialogues include assistant hallucinations (unconfirmed details), implicit expressions, and coreference. Two versions:

HaluMem-Medium : 20 users, 30,073 dialogue turns, ~160k tokens/user, 14,948 memory points, 3,467 QA pairs.

HaluMem-Long : Same core memories, but irrelevant dialogue inserted within and between sessions, expanding context to ~1M tokens/user, 53,516 total turns. The Long version adds distraction, not new facts, enabling study of retention under noise.

QA covers six types: basic recall, multi-hop reasoning, dynamic update, memory boundary, generalization, conflict detection. Human evaluation of 700 Medium sessions achieved 95.70% correctness.

Figure 2: HaluMem's six-stage construction pipeline. From user profile, life skeleton, event stream to link memory points with dialogues and questions. (Paper Figure 3)
Figure 2: HaluMem's six-stage construction pipeline. From user profile, life skeleton, event stream to link memory points with dialogues and questions. (Paper Figure 3)

Experiments: Six Memory Systems Tested

Evaluated systems: Mem0, Mem0-Graph, Memobase, MemOS, Supermemory, Zep. Unified settings: top-10 memories for update verification, top-20 for QA, GPT-4o for answer generation. Zep omitted from extraction metrics due to API limits.

Table 1: Results of six systems on memory extraction, update, and QA tasks. Multiple metrics together distinguish coverage, correctness, omission. (Paper Table 3)
Table 1: Results of six systems on memory extraction, update, and QA tasks. Multiple metrics together distinguish coverage, correctness, omission. (Paper Table 3)

Finding 1: More Memories ≠ More Reliable Memories

On Medium, extraction recall: Mem0 42.91%, Mem0-Graph 43.28%, Memobase 14.55%. On Long, recall drops drastically: 3.23%, 2.24%, 6.18% — core memories unchanged, only noise added.

Conversely, MemOS (81.90% recall on Long) and Supermemory (53.02%) cover more targets but suffer low False Memory Resistance (FMR): 28.85% and 36.86% respectively. FMR measures ability to ignore assistant-invented, unconfirmed details. High recall does not automatically confer strong false-memory filtering.

Reliable memory requires both: don't miss what should be recorded, don't store what shouldn't be trusted.

Finding 2: Low Update Hallucination Rate Masks Massive Omissions

Update hallucination rate (errors among attempted updates) is <1.2% for all systems on both versions. However, update omission rates on Long exceed 60% for five of six systems. Mem0 and Mem0-Graph omission rates: 98.51% and 98.40%, yielding correct update rates of only 1.45% and 1.47%. MemOS leads with 65.25% correct updates but still misses ~1/3.

Why low hallucination coexists with high omission? Because hallucination rate is computed only over attempted updates; if few updates are attempted, the error proportion stays low. Moreover, updates depend on prior extraction: if a fact was never stored, it cannot be updated. Operation-level evaluation reveals that update failures often originate at the first write.

Finding 3: Upstream Memory Errors Propagate to Final Answers

Under unified QA, no system reaches 70% accuracy. MemOS tops at 67.23% (Medium) and 64.44% (Long). Mem0 drops from 53.02% to 28.11% accuracy, while its omission rate rises from 27.81% to 54.60%, mirroring its extraction recall collapse.

Question-type analysis: systems handle unknown-info and false-premise detection relatively well, but struggle with multi-hop reasoning, dynamic updates, and generalization. Final answers reflect the entire memory pipeline; improving them requires fixing extraction, update, and retrieval jointly.

Robustness check: re-evaluating 10 users with GPT-5.4-mini as judge changes absolute scores but preserves rankings, trade-offs, and core findings.

Figure 3: Six memory systems' answer accuracy on six question types, showing performance on both evaluation sets. Different types test different memory usage scenarios and long-context impact. (Paper Figure 5)
Figure 3: Six memory systems' answer accuracy on six question types, showing performance on both evaluation sets. Different types test different memory usage scenarios and long-context impact. (Paper Figure 5)

From "Can Remember" to "Remembers Reliably"

HaluMem makes improvement directions concrete: if extraction fails, check coverage and source credibility; if update fails, verify old-new linking and completion; if QA fails, trace back through retrieved memories and upstream records. Single metrics are misleading: high recall may admit false memories, low update hallucination may hide omissions, and final QA scores can mask internal failures.

Limitations: current data focuses on profile/event/relation memories, not tool/skill memories; system API openness affects observability. Nevertheless, HaluMem provides a reusable methodology: verify each memory operation behind the final answer, turning "where did it go wrong?" into a measurable question.

Returning to the opening example: next time the assistant gets the job title wrong, we can ask: was it recorded correctly initially? Was it updated after promotion? Which record was used at answer time? When these questions are answerable, AI moves closer to truly reliable long-term memory.

Paper: https://arxiv.org/pdf/2511.03506 Open-source:

https://github.com/MemTensor/HaluMem
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

memory extractionEMNLP 2026AI agent memoryhallucination benchmarkHaluMemmemory updateMemTensor
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.