Cross-Session Memory Poisoning: Auditable Secure Write Paths for Agent Memory

This article explains how memory poisoning attacks persist across AI agent sessions, details the MINJA attack technique, and proposes auditable memory storage formats, secure write paths with trust scoring, and detection methods including integrity checks and activation monitoring to defend against persistent memory corruption.

DeepHub IMBA
DeepHub IMBA
DeepHub IMBA
Cross-Session Memory Poisoning: Auditable Secure Write Paths for Agent Memory

Memory Poisoning Enters OWASP Top 10

Memory poisoning (ASI06 in the OWASP Agentic Applications Top 10) achieves over 95% injection success rates using ordinary queries. Most teams evaluate agent memory on recall metrics, chunk size, rerankers, and embedding models — all focused on "how to read" — but rarely scrutinize how records enter storage in the first place. When an agent processes a web page, PDF, support ticket, or another agent's output, it decides what to persist without leaving an audit trail: source document, human review status, or whether the content is fact versus instruction.

The core problem: retrieval faithfully returns a record that should never have been written.

Poisoned Records Outlive Their Creating Session

Traditional prompt injection is confined to a single context window; memory removes that boundary, changing the attack structure itself. An attacker plants instructions in an untrusted document; the agent persists a fragment. Days or weeks later, a different user in a different session triggers retrieval, and the payload executes — the attacker need not be present. Existing defenses fail because injection detection watches content entering the context, but at trigger time the malicious content comes from the system's own trusted store. Session isolation does not help because the attack targets persistence itself.

Modern techniques go beyond simple injection strings. MINJA adds bridging steps that mimic normal reasoning, then uses indication prompts to make the agent generate a "worth remembering" version of the payload, gradually stripping recognizable attack features while retaining the poisoning effect. The final stored text looks benign because it was rewritten in the agent's own words. Another pattern is delayed triggering, e.g., "if the user later says yes, apply this update"; a future unrelated "sure" may be interpreted as authorization.

Two amplifying consequences: a compromised agent may rationalize poisoned memory as learned knowledge, and in multi-agent systems the corrupted memory propagates through inter-agent communication into stores that never saw the original document.

Auditable Memory Records

Treating memory as data rather than a service has led to file formats with built-in security properties. memoryfield is a flat directory of Markdown pages with YAML frontmatter, distributed as a zip. Filenames use lowercase ASCII, UTF-8 encoding; each page ≤ 8,192 bytes (~1,300 words). Example frontmatter:

---  
title: Carbon Fibre Woks  
created: '2026-03-01T09:00:00Z'  
updated: '2026-08-22T14:30:00Z'  
uuid: 6aa615f0-486f-48a7-a210-ba4f5ff18c8b  
summary: Thermal properties of carbon fibre cookware  
---  
Carbon fibre woks conduct heat evenly, but...

A SQLite index file sits beside the pages, named after the embedding model used:

CREATE TABLE pages (  
    filename      TEXT PRIMARY KEY,  
    frontmatter   JSON NOT NULL,  
    last_modified DATETIME NOT NULL,  
    sha256_hash   BLOB NOT NULL,  
    embedding     BLOB NOT NULL  
);

Benefits: every record is a plain file readable without client libraries; the whole store can be diffed in git; each page has a hash so tampering is detectable; full prose is preserved (not chunked), so human review sees complete contextual paragraphs.

Though not designed as security features, these properties directly serve the threat model for detecting memory poisoning. The mechanism is simple — agents can use grep, SQLite, and bash — and survives model updates without vendor lock-in.

Reliable Write Path

Auditability alone doesn't stop bad writes; the entry point needs a gate. Every memory write must be treated as untrusted input crossing a privilege boundary:

def commit_memory(text, source, session, author):  
    if looks_like_instruction(text):  
        return quarantine(text, reason="imperative_content")  
    return write_page(text, frontmatter={  
        "source_uri": source,  
        "session_id": session,  
        "authored_by": author,  
        "review_state": "unreviewed",  
        "trust": 0.3 if author != "human" else 1.0,  
        "expires": iso(now() + timedelta(days=90))  
    })

Imperative content must be stripped before persistence, not at retrieval. Memory should record "X is true", not "do Y"; imperative mood is a strong risk signal. Source metadata (document, session, author, timestamp) must be captured at write time — retroactive addition is usually impossible. Trust is not binary: unreviewed memory is weak evidence ("Agent reported X on date"), while human-reviewed content carries higher authority. If AI modifies an approved record, review state must reset; otherwise the original approval "washes" onto the new content.

Retrieval must respect trust differences: low-trust records are down-weighted even with high similarity; old memories decay over time; anomalous activation patterns are monitored. Memory TTL should extend on successful recall, not on retrieval frequency, lest "often retrieved" become a proxy for "correct".

Detecting Whether a System Is Already Poisoned

Most production systems already have memory stores, so the task is improvement, not greenfield design. Existing records lacking provenance can be prioritized by suspicion. First, find records with no human or document explanation:

SELECT filename, json_extract(frontmatter,'$.source_uri') AS src,  
       json_extract(frontmatter,'$.review_state') AS review, last_modified  
FROM pages  
WHERE src IS NULL  
   OR review IS NULL  
   OR lower(json_extract(frontmatter,'$.summary')) GLOB '*[[]if the user*'  
ORDER BY last_modified DESC;

Then verify integrity via the sha256_hash column: recompute each page's hash and compare with the index. In a store where only the indexer writes, mismatches should be zero; any non-zero result warrants investigation.

Content inspection should not only hunt known payloads but also grep for attack-typical phrasing: "if the user", "when asked about", "from now on", "always respond", "remember to". Normal memories are declarative ("auth service uses rotating JWT"), not conditional instructions. Imperative tone doesn't equal malice but flags the record for review.

Retrieval behavior analysis: establish a baseline of which memories are normally recalled. If a long-dormant record suddenly activates across unrelated sessions and users, alert. Delayed-trigger attacks rely on this pattern — the record sits idle until a common word appears.

Two architectural imperatives: (1) write logs must be immutable and stored outside the memory store; a compromised agent that can alter audit trails defeats auditing. (2) A circuit breaker is needed to freeze memory reads for a specific agent or tenant without shutting down the agent entirely, avoiding the choice between running polluted or stopping completely.

Cost Considerations

The file format has explicit limits: flat directories become unwieldy at scale; 8 KB page caps split complex topics; a single 768-dimension embedding per page is coarser than tuned chunk-level hybrid retrieval; and irrelevant memories accumulate, requiring periodic cleanup. Chunk-level retrieval answers some questions page-level indexes miss, but similarity is probabilistic — only adding provenance makes results truly useful. Time-sensitive queries degrade with similarity search, often returning both obsolete and current decisions together. Moderation and sanitization thresholds need careful calibration (precision/recall trade-off) — they cannot be set once and forgotten.

If an agent has no persistent memory, these issues don't apply. Single-user, local-only memory that never ingests third-party content doesn't urgently need source tagging. The concern is shared memory scenarios — cross-user, cross-agent, and any ingestion from the open web.

Summary

Agent memory changes the nature of "storing a record". A conventional database record is an application-written fact; an agent memory record is a model-decided sentence, provenance unrecorded, later read by another model as instruction-shaped context. They are not the same object, so control mechanisms and security requirements must differ.

Author: Sebastian Buzdugan

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agent securitymemory poisoningauditable memorymemoryfield formatMINJA attackOWASP ASI06secure write pathtrust scoring
DeepHub IMBA
Written by

DeepHub IMBA

A must‑follow public account sharing practical AI insights. Follow now. internet + machine learning + big data + architecture = IMBA

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.