Agent Memory Leaderboard Launched: The First Open Benchmark for Long‑Term Memory Systems

The Agent Memory Leaderboard (AML) debuted on July 29, 2026, offering a unified, reproducible evaluation framework that combines multi‑source text and code memory datasets, standardized protocols, ability profiling, and low‑barrier integration to fairly compare memory systems while providing detailed performance diagnostics and incentives for participants.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Agent Memory Leaderboard Launched: The First Open Benchmark for Long‑Term Memory Systems

Unified Evaluation: Multi‑Source Data and Full Process

AML integrates over ten mainstream text‑memory datasets (PersonaMem, LoCoMo‑Refined, CLBench, BEAM, LongMemEval, ScriptMem, etc.), covering more than 1,500 dialogues, approximately 1.5 billion characters and around 5,000 evaluation questions. It also includes a code‑memory benchmark constructed from 12 GitHub repositories, 150 base tasks and 1,290 finely annotated historical tasks. The platform enforces a unified Add API for writing memory, a unified Search API for retrieval, a fixed Answer model for response generation, and a standardized Eval pipeline for scoring, so that score differences primarily reflect the memory system itself.

Unified evaluation diagram
Unified evaluation diagram

Ability Profiling: Reconstructing Capability Dimensions

Tasks from diverse datasets are remapped onto a unified set of capability dimensions: explicit fact recall, multi‑hop reasoning, temporal/event sequencing, memory governance, personalization, rule execution, epistemology, and safety/privacy. This profiling produces per‑dimension performance scores, revealing concrete strengths and weaknesses of each system.

Ability profile illustration
Ability profile illustration

Low‑Barrier Integration: Supporting Long‑Term Operation

Participants need only implement two core interfaces— Add for memory write and Search for retrieval. Answer generation, scoring, and result aggregation are performed centrally, enabling both open‑source methods (submitted with code, configuration and reproducibility materials) and commercial products (accessed via stable APIs) to be benchmarked continuously.

First Agent Memory Challenge 2026

The inaugural competition launched on 2026‑07‑29. It provides two leaderboard tracks: an Open‑Source Methods track (rankings tied to code and reproducibility) and a Commercial Products track (rankings based on stable API offerings). Participants follow a four‑step workflow: registration, API integration, test submission, and result verification.

Key dates: registration opens 2026‑07‑29, submission deadline 2026‑08‑07, results announced mid‑August.

Official repository (protocol, submission guide, integration examples): https://github.com/AML-memory/agent-memory-leaderboard

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AIbenchmarkEvaluationAgent MemoryLong-Term MemoryOpen Leaderboard
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.