Cross-Session Memory Poisoning: Auditable Secure Write Paths for Agent Memory
This article explains how memory poisoning attacks persist across AI agent sessions, details the MINJA attack technique, and proposes auditable memory storage formats, secure write paths with trust scoring, and detection methods including integrity checks and activation monitoring to defend against persistent memory corruption.
Memory Poisoning Enters OWASP Top 10
Memory poisoning (ASI06 in the OWASP Agentic Applications Top 10) achieves over 95% injection success rates using ordinary queries. Most teams evaluate agent memory on recall metrics, chunk size, rerankers, and embedding models — all focused on "how to read" — but rarely scrutinize how records enter storage in the first place. When an agent processes a web page, PDF, support ticket, or another agent's output, it decides what to persist without leaving an audit trail: source document, human review status, or whether the content is fact versus instruction.
The core problem: retrieval faithfully returns a record that should never have been written.
Poisoned Records Outlive Their Creating Session
Traditional prompt injection is confined to a single context window; memory removes that boundary, changing the attack structure itself. An attacker plants instructions in an untrusted document; the agent persists a fragment. Days or weeks later, a different user in a different session triggers retrieval, and the payload executes — the attacker need not be present. Existing defenses fail because injection detection watches content entering the context, but at trigger time the malicious content comes from the system's own trusted store. Session isolation does not help because the attack targets persistence itself.
Modern techniques go beyond simple injection strings. MINJA adds bridging steps that mimic normal reasoning, then uses indication prompts to make the agent generate a "worth remembering" version of the payload, gradually stripping recognizable attack features while retaining the poisoning effect. The final stored text looks benign because it was rewritten in the agent's own words. Another pattern is delayed triggering, e.g., "if the user later says yes, apply this update"; a future unrelated "sure" may be interpreted as authorization.
Two amplifying consequences: a compromised agent may rationalize poisoned memory as learned knowledge, and in multi-agent systems the corrupted memory propagates through inter-agent communication into stores that never saw the original document.
Auditable Memory Records
Treating memory as data rather than a service has led to file formats with built-in security properties. memoryfield is a flat directory of Markdown pages with YAML frontmatter, distributed as a zip. Filenames use lowercase ASCII, UTF-8 encoding; each page ≤ 8,192 bytes (~1,300 words). Example frontmatter:
---
title: Carbon Fibre Woks
created: '2026-03-01T09:00:00Z'
updated: '2026-08-22T14:30:00Z'
uuid: 6aa615f0-486f-48a7-a210-ba4f5ff18c8b
summary: Thermal properties of carbon fibre cookware
---
Carbon fibre woks conduct heat evenly, but...A SQLite index file sits beside the pages, named after the embedding model used:
CREATE TABLE pages (
filename TEXT PRIMARY KEY,
frontmatter JSON NOT NULL,
last_modified DATETIME NOT NULL,
sha256_hash BLOB NOT NULL,
embedding BLOB NOT NULL
);Benefits: every record is a plain file readable without client libraries; the whole store can be diffed in git; each page has a hash so tampering is detectable; full prose is preserved (not chunked), so human review sees complete contextual paragraphs.
Though not designed as security features, these properties directly serve the threat model for detecting memory poisoning. The mechanism is simple — agents can use grep, SQLite, and bash — and survives model updates without vendor lock-in.
Reliable Write Path
Auditability alone doesn't stop bad writes; the entry point needs a gate. Every memory write must be treated as untrusted input crossing a privilege boundary:
def commit_memory(text, source, session, author):
if looks_like_instruction(text):
return quarantine(text, reason="imperative_content")
return write_page(text, frontmatter={
"source_uri": source,
"session_id": session,
"authored_by": author,
"review_state": "unreviewed",
"trust": 0.3 if author != "human" else 1.0,
"expires": iso(now() + timedelta(days=90))
})Imperative content must be stripped before persistence, not at retrieval. Memory should record "X is true", not "do Y"; imperative mood is a strong risk signal. Source metadata (document, session, author, timestamp) must be captured at write time — retroactive addition is usually impossible. Trust is not binary: unreviewed memory is weak evidence ("Agent reported X on date"), while human-reviewed content carries higher authority. If AI modifies an approved record, review state must reset; otherwise the original approval "washes" onto the new content.
Retrieval must respect trust differences: low-trust records are down-weighted even with high similarity; old memories decay over time; anomalous activation patterns are monitored. Memory TTL should extend on successful recall, not on retrieval frequency, lest "often retrieved" become a proxy for "correct".
Detecting Whether a System Is Already Poisoned
Most production systems already have memory stores, so the task is improvement, not greenfield design. Existing records lacking provenance can be prioritized by suspicion. First, find records with no human or document explanation:
SELECT filename, json_extract(frontmatter,'$.source_uri') AS src,
json_extract(frontmatter,'$.review_state') AS review, last_modified
FROM pages
WHERE src IS NULL
OR review IS NULL
OR lower(json_extract(frontmatter,'$.summary')) GLOB '*[[]if the user*'
ORDER BY last_modified DESC;Then verify integrity via the sha256_hash column: recompute each page's hash and compare with the index. In a store where only the indexer writes, mismatches should be zero; any non-zero result warrants investigation.
Content inspection should not only hunt known payloads but also grep for attack-typical phrasing: "if the user", "when asked about", "from now on", "always respond", "remember to". Normal memories are declarative ("auth service uses rotating JWT"), not conditional instructions. Imperative tone doesn't equal malice but flags the record for review.
Retrieval behavior analysis: establish a baseline of which memories are normally recalled. If a long-dormant record suddenly activates across unrelated sessions and users, alert. Delayed-trigger attacks rely on this pattern — the record sits idle until a common word appears.
Two architectural imperatives: (1) write logs must be immutable and stored outside the memory store; a compromised agent that can alter audit trails defeats auditing. (2) A circuit breaker is needed to freeze memory reads for a specific agent or tenant without shutting down the agent entirely, avoiding the choice between running polluted or stopping completely.
Cost Considerations
The file format has explicit limits: flat directories become unwieldy at scale; 8 KB page caps split complex topics; a single 768-dimension embedding per page is coarser than tuned chunk-level hybrid retrieval; and irrelevant memories accumulate, requiring periodic cleanup. Chunk-level retrieval answers some questions page-level indexes miss, but similarity is probabilistic — only adding provenance makes results truly useful. Time-sensitive queries degrade with similarity search, often returning both obsolete and current decisions together. Moderation and sanitization thresholds need careful calibration (precision/recall trade-off) — they cannot be set once and forgotten.
If an agent has no persistent memory, these issues don't apply. Single-user, local-only memory that never ingests third-party content doesn't urgently need source tagging. The concern is shared memory scenarios — cross-user, cross-agent, and any ingestion from the open web.
Summary
Agent memory changes the nature of "storing a record". A conventional database record is an application-written fact; an agent memory record is a model-decided sentence, provenance unrecorded, later read by another model as instruction-shaped context. They are not the same object, so control mechanisms and security requirements must differ.
Author: Sebastian Buzdugan
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DeepHub IMBA
A must‑follow public account sharing practical AI insights. Follow now. internet + machine learning + big data + architecture = IMBA
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
