Why AI Agents Need Documentation, Not Memory: A Critique of RAG-Based Memory Plugins
The article critiques typical RAG-based memory plugins for AI agents, highlighting five fundamental flaws—similarity retrieval ignores correctness, fragments lose context, outdated facts persist, agents cannot recognize knowledge gaps, and storage is unauditable—and proposes a documentation-centric approach where agents read and write Markdown files in a structured workspace, exemplified by the open-source Operator Memory plugin.
The article opens by describing the typical workflow of AI memory plugins: they analyze conversation history, generate thousands of fragments, store them in a vector database, and retrieve the top‑5 most similar fragments for each prompt. If that fails, the agent is given a search tool to query the database itself. The author notes that almost every memory plugin follows this same pattern, sometimes adding layers like long‑term/short‑term separation, background deduplication daemons, "dreamer" processes that rewrite memories at night, or rerankers—but the core architecture remains flawed.
Five Hard Flaws of RAG‑Style Memory
The author identifies five fundamental problems that all such plugins share:
Similarity retrieval only cares about distance, not correctness. Vector search returns embeddings that are close in space, with no regard for which fragment is accurate, up‑to‑date, or missing critical information.
Fragments are stripped of context. A single RAG chunk can hold only limited information; the motivation, environment, constraints, and lessons learned at the time are lost.
The past is treated as fact. Code changes daily. Hundreds of fragments about authentication may be obsolete, yet there is no mechanism to know which ones are still valid.
The agent does not know what it does not know. Even with a search tool, the agent cannot know which queries to issue; unknown unknowns remain invisible.
Storage is unauditable. Ten thousand embeddings sit in a SQLite database. There is no way to tell which are stale, which have never been retrieved, or which are silently biasing the agent's decisions.
These issues stem from a single root cause: the assumption that the agent forgets things, so we must help it remember more. The author argues this is not how humans handle knowledge—people write things down and then consult the records, rather than re‑watching years of meeting recordings.
A Different Approach: Documentation, Not Memory
In software projects, teams maintain READMEs, specs, decision records, and research notes. They consult these documents before working and update them afterward. Agents can operate the same way. Give the agent a readable, writable Markdown workspace containing instructions, specifications, decisions, indexes, and research notes. The work loop shifts from prompt → build → forget to prompt → consult → build → update .
No vector database, no embeddings, no background processes are required. Every document is directly viewable, editable, committable, and shareable with the team.
Operator Memory: An Open‑Source Implementation
The author has used this approach for over a year and released the Operator Memory plugin (https://github.com/aerovato/operator-memory), compatible with Claude Code, Codex, OpenCode, and other platforms. Installation is a single npm command; after initialization the agent automatically maintains documentation in the workspace.
Knowledge is organized in three layers: .operator/ – private project‑specific knowledge .operator-shared/ – knowledge that can be pushed with the repository ~/.operator/user/ – personal rules shared across projects
Each session follows the same cycle: consult existing knowledge, do the work, update the documentation.
Counterarguments and Discussion
Game developer Mario Zechner (creator of libGDX) disagrees:
Recommended reading, but I still disagree. For code, neither memory nor documentation is efficient. Your codebase is all you need. Make it modular, keep modules small, maintain at most a small map to navigate.
Some commenters support this view: "The codebase is the map; the agent keeps asking where the map is." Others object: "Documentation and comments go stale quickly and can seriously mislead." Mario replies to a comment about agents being bad at writing modular code: "If you explicitly tell them how, they actually can."
A more nuanced angle emerges: "Most knowledge work has no codebase at all; traces are scattered across a dozen websites, so memory plugins end up doing the job a repository should do." This highlights the key distinction: for pure code projects, a minimal map may suffice. But for user preferences, product decisions, cross‑system dependencies, and research conclusions—things that cannot be expressed in code—a maintainable documentation space is far more reliable than scraping conversation history.
Memory plugins essentially fight forgetting. Documentation does not; it is written for the future, capturing only the conclusions worth keeping at this moment. With documentation, you can still review what the agent knows. With ten thousand embeddings, you cannot even locate the errors.
The debate has no single answer, but it forces a crucial question: what do you actually want the agent to remember—past chat logs, or the current facts of the project?
Original article: https://liao.gg/blog/agents-dont-need-memory GitHub repository: https://github.com/aerovato/operator-memory
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
AI Engineering
Focused on cutting‑edge product and technology information and practical experience sharing in the AI field (large models, MLOps/LLMOps, AI application development, AI infrastructure).
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
