Supermemory: Open-Source Memory Layer for AI Agents Hits 95% Recall at 99.4% Compression
The article analyzes supermemory, an open-source memory engine that gives AI agents persistent, stateful memory across sessions, achieving 95% recall on LongMemEval while adding only ~720 tokens, and provides Spring AI integration code and self-hosted deployment guidance.
Problem: Context Window Limits and Token Costs
When chat history exceeds the model's context window, older messages are truncated, causing agents to forget facts like user allergies or order numbers. Stuffing full history into the prompt makes input tokens grow quadratically with each turn, inflating costs and latency. Truncation loses facts, not just characters.
Solution: supermemory as Independent Memory Infrastructure
supermemory (GitHub: supermemoryai/supermemory, 30k stars) is a memory and context engine delivered via API (MIT license). Agents offload conversations and documents to it; the engine extracts facts, maintains user profiles, handles contradictions and expiration, and returns relevant context for the next turn.
Architecture: a core memory engine handles fact extraction, update tracking, and automatic forgetting. On top sit four modules: user profiles, hybrid search, connectors (Google Drive, Gmail, Notion, GitHub), and file processing (PDF, images, video, code).
Key Technical Distinction: Memory vs. RAG
RAG is stateless retrieval — same document returns same chunks for every user. Memory is stateful: it tracks how facts about a specific user change over time. If a user moves from New York to San Francisco, the system understands the later statement overrides the earlier one.
supermemory maintains per-user profiles split into static (stable facts: "senior engineer", "uses Vim") and dynamic (recent context: "debugging rate-limiting this week"). Fetching a profile takes ~50 ms and can be injected into the system prompt. Transient facts ("I have an exam tomorrow") auto-expire after the date. Contradictory statements are resolved by the engine without custom cleaning logic.
Benchmarks and Caveats
Official claims: first place on LongMemEval, LoCoMo, and ConvoMem benchmarks. LongMemEval Recall@15 reaches 95% with only ~720 tokens added (99.4% compression). On the xAFS benchmark (110 questions), Claude token usage dropped from 72M to 24M; Codex saved 1.75x.
However, these numbers come mainly from official evaluations. Third-party replication on LoCoMo scored ~70% with limited samples. The team open-sourced MemoryBench, a framework that can benchmark Mem0, Zep, and supermemory side-by-side — a credible move. The author advises running MemoryBench on your own dialogue data before adopting.
Spring AI Integration (Java)
Official SDKs exist only for TypeScript and Python. The REST API surface is three endpoints, easily wrapped with Spring's RestClient:
Write: POST /v3/documents
Search: POST /v4/search
Profile: POST /v4/profile
@Bean
RestClient supermemory(RestClient.Builder builder,
@Value("${supermemory.api-key}") String apiKey,
@Value("${supermemory.base-url:https://api.supermemory.ai}") String baseUrl) {
return builder.baseUrl(baseUrl)
.defaultHeader(HttpHeaders.AUTHORIZATION, "Bearer " + apiKey)
.build();
}Store each conversation turn:
public void remember(String userId, String question, String answer) {
restClient.post().uri("/v3/documents")
.body(Map.of(
"content", "Q: " + question + "
A: " + answer,
"containerTag", userId,
"metadata", Map.of("source", "customer-service")
))
.retrieve()
.toBodilessEntity();
}Before the next turn, fetch the profile and call Spring AI's ChatClient:
public String chat(String userId, String question) {
JsonNode res = restClient.post().uri("/v4/profile")
.body(Map.of("containerTag", userId, "q", question))
.retrieve()
.body(JsonNode.class);
String staticFacts = res.at("/profile/static").toString();
String dynamicFacts = res.at("/profile/dynamic").toString();
String reply = chatClient.prompt()
.system("用户长期画像:" + staticFacts + "
用户近期动态:" + dynamicFacts)
.user(question)
.call()
.content();
remember(userId, question, reply);
return reply;
}Three integration pitfalls: containerTag is singular, placed in the request body, and must be consistent for reads and writes — it is the multi-tenant isolation key. Many outdated tutorials use plural containerTags and the deprecated /v3/search endpoint.
Writes are asynchronous; searching before the document state reaches done returns empty results. For immediate visibility, include dreaming: "instant" on write.
The code snippets show the integration skeleton; refer to official docs for exact field contracts.
For non-Spring environments, official plugins exist for Claude Code, Cursor, Codex, OpenCode, and an MCP server at https://mcp.supermemory.ai/mcp enabling cross-session memory in IDE agents.
Data Storage: Cloud vs. Self-Hosted
Two deployment paths:
Managed cloud (console.supermemory.ai): data stored in Postgres + Cloudflare Workers. Free tier to $399/month. Convenient but raises data egress and compliance concerns because memories contain user PII.
Self-hosted : single binary runs the full engine. Install via curl -fsSL https://supermemory.ai/install | bash or npx supermemory local. API serves on localhost:6767. Embeddings default to local Xenova/bge-base-en-v1.5; data persists in ./.supermemory. For fully offline use, plug in Ollama (e.g., gpt-oss:20b). Switching from cloud to local only requires changing the base URL; the API is identical.
Author's preference: use cloud for PoC, switch to self-hosted before production — keeping memory data on-premises is safer.
A real incident (GitHub issue #792) illustrates operational risk: after an org/account merge, a user's 3000+ documents became unreadable despite successful writes. The fix was applied, but it underscores that the memory layer is a stateful service — upgrades and migrations must be treated like database changes, with regression tests verifying "write then read" correctness.
Memory failures are silent: HTTP 200 but wrong facts or wrong user. Debug by calling /v3/documents and /v4/profile directly to isolate layers.
Open Issues
The open-source repo contains SDKs, MCP server, IDE plugins, and web app shell — the core memory engine is not in the repo; it ships as a managed service and a single binary. Source-level audit of the engine is not possible.
Quality hinges on extraction accuracy. If the engine mis-extracts facts or stores noise, downstream profiles and search inherit the errors silently — agents confidently hallucinate. Pre-production testing with real dialogue samples to measure extraction precision is mandatory.
Author's Verdict
Larger context windows don't cure "amnesia per conversation" because the root cause is lack of a durable fact store. supermemory turns that layer into standalone infrastructure. Java integration takes dozens of HTTP lines; self-hosting offers a data-locality escape hatch. Don't chase star counts — run MemoryBench on your own data and let accuracy decide.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architecture Digest
Focusing on Java backend development, covering application architecture from top-tier internet companies (high availability, high performance, high stability), big data, machine learning, Java architecture, and other popular fields.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
