Memory Poisoning Across Sessions: Building Auditable Secure Write Paths for Agent Memory

The article explains how memory poisoning attacks persist across sessions in AI agents, details the memoryfield format for auditable storage, proposes a secure write path with trust levels and source tracking, and provides detection queries for identifying poisoned memories.

Data Party THU
Data Party THU
Data Party THU
Memory Poisoning Across Sessions: Building Auditable Secure Write Paths for Agent Memory

Memory poisoning (ASI06 in OWASP Agentic Applications Top 10) achieves over 95% injection success with ordinary queries. Unlike prompt injection confined to a single context window, poisoned memories persist for days or weeks and execute when retrieved by unrelated users in future sessions.

Why Poisoned Records Outlive Their Creating Session

Attackers plant instructions in untrusted documents. The agent processes the document and persists a fragment. Later, a different user in a different session triggers retrieval, and the payload executes. The attacker need not be present. Existing defenses fail because injection detection watches content entering the context window; by retrieval time, the malicious content comes from the system's own trusted storage. Session isolation does not help because the attack targets persistence itself.

Modern techniques like MINJA add bridging steps that look like normal reasoning, then use indication prompts to make the agent rewrite the payload in its own words, stripping obvious attack markers while retaining the poisoning effect. Delayed triggers (e.g., "if the user later says yes, apply this update") exploit common words like "sure" as authorization. In multi-agent systems, corrupted memories propagate through inter-agent communication into stores that never saw the original document.

Auditable Memory Records: The memoryfield Format

Treating memory as data rather than a service enables auditability. The memoryfield format uses a flat directory of Markdown pages with YAML frontmatter, distributed as a zip. Filenames are lowercase ASCII, UTF-8 encoded; each page ≤ 8,192 bytes (~1,300 words). Example frontmatter:

---  
title: Carbon Fibre Woks  
created: '2026-03-01T09:00:00Z'  
updated: '2026-08-22T14:30:00Z'  
uuid: 6aa615f0-486f-48a7-a210-ba4f5ff18c8b  
summary: Thermal properties of carbon fibre cookware  
---  
Carbon fibre woks conduct heat evenly, but...

A SQLite index file sits alongside pages, named after the embedding model used:

CREATE TABLE pages (  
    filename      TEXT PRIMARY KEY,  
    frontmatter   JSON NOT NULL,  
    last_modified DATETIME NOT NULL,  
    sha256_hash   BLOB NOT NULL,  
    embedding     BLOB NOT NULL  
);

Benefits: each record is a plain file readable without client libraries; the whole store can be diffed in git; every page has a hash so tampering is detectable; full prose is preserved (not chunked), so human review sees complete paragraphs in context. These properties, though not designed as security features, directly address the threat model. The mechanism is simple—agents can use grep, SQLite, and bash—and survives model updates without vendor lock-in.

Reliable Write Path

Auditability alone does not stop bad writes. Every memory write must treat untrusted input as crossing a trust boundary:

def commit_memory(text, source, session, author):  
    if looks_like_instruction(text):  
        return quarantine(text, reason="imperative_content")  
    return write_page(text, frontmatter={  
        "source_uri": source,  
        "session_id": session,  
        "authored_by": author,  
        "review_state": "unreviewed",  
        "trust": 0.3 if author != "human" else 1.0,  
        "expires": iso(now() + timedelta(days=90))  
    })

Imperative content must be stripped before persistence, not at retrieval. Memory should record "X is true," not "do Y." Imperative language in a record is a strong risk signal. Source metadata (document, session, author, timestamp) must be captured at write time; retrofitting is usually impossible. Trust is not binary: unreviewed memories are weak evidence ("an agent reported this on date D"), while human-reviewed content carries higher authority. If an AI modifies an approved memory, the review state must reset; otherwise the original approval "washes" onto the new content.

Retrieval must respect trust differences: low-trust records are down-weighted even with high similarity; old memories decay over time; anomalous activation patterns are monitored. TTL should extend on successful recall, not on retrieval frequency, to avoid conflating "often retrieved" with "correct."

Detecting Whether a System Is Already Poisoned

Most production systems already have memory stores, so the task is improvement, not greenfield design. Prioritize suspicious records: those lacking human or document provenance. Example query:

SELECT filename, json_extract(frontmatter,'$.source_uri') AS src,  
       json_extract(frontmatter,'$.review_state') AS review, last_modified  
FROM pages  
WHERE src IS NULL  
   OR review IS NULL  
   OR lower(json_extract(frontmatter,'$.summary')) GLOB '*[[]if the user*'  
ORDER BY last_modified DESC;

Then verify integrity by recomputing sha256_hash for each page and comparing with the index. In a store where only the indexer writes, mismatches should be zero; any non-zero result warrants investigation. Content scanning with grep for attack patterns ("if the user", "when asked about", "from now on", "always respond", "remember to") flags records for review—normal memories are declarative, not imperative. Finally, monitor retrieval behavior: establish a baseline of which memories are normally retrieved; a long-dormant record suddenly activating across unrelated sessions and users signals a delayed trigger.

Two architectural requirements: (1) write logs must be immutable and stored outside the memory store; a compromised agent that can alter audit trails defeats auditing. (2) A circuit breaker must freeze memory reads for a specific agent or tenant without shutting down the agent entirely, avoiding the choice between running polluted or stopping completely.

Cost Considerations

The file format has limits: flat directories become unwieldy at scale; 8 KB pages split complex topics; a single 768-dimension embedding per page is coarser than tuned chunk-level hybrid retrieval; irrelevant memories accumulate and require periodic cleanup. Chunk-level retrieval answers some questions page-level indexes cannot, but similarity is probabilistic—only adding provenance makes results truly useful. Time-based queries degrade with similarity search, often returning both superseded and current decisions together. Moderation and sanitization thresholds need careful calibration (precision/recall trade-off) or they either filter legitimate memories or miss stealthy attacks. If an agent has no persistent memory, or memory is single-user, local, and never ingests third-party content, these measures are not urgent. They matter for shared memory—cross-user, cross-agent, and any system ingesting open-web input.

Summary

Agent memory changes the nature of "storing a record." A conventional database record is a fact written by application code; an agent memory record is a sentence a model decided to keep, with no recorded provenance, later read by another model as instruction-shaped context. They are not the same object, so control mechanisms and security requirements must differ.

Author: Sebastian Buzdugan

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI securityagent memoryaudit trailmemory poisoningMINJA attackOWASP ASI06secure write pathmemoryfield
Data Party THU
Written by

Data Party THU

Official platform of Tsinghua Big Data Research Center, sharing the team's latest research, teaching updates, and big data news.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.