WeChat's Agentic OLAP Architecture: Observability, Memory & Access Control

WeChat's technical architecture team shares their exploration of adapting OLAP infrastructure for Agentic AI, covering distributed observability with Langfuse and ClickHouse, remote memory services using vector search and FUSE, and access protection mechanisms treating agents as authenticated identities.

DataFunTalk
DataFunTalk
DataFunTalk
WeChat's Agentic OLAP Architecture: Observability, Memory & Access Control

OLAP Meets Agentic: From Serving Humans to Serving Agents

Agentic AI is reshaping software interaction from human-operated tools to autonomous decision-making agents with perception, memory, reasoning, and action capabilities. WeChat identifies 2026 as the "Agent元年" (Agent元年) where model maturity and engineering readiness converge. Common WeChat agent scenarios include platform customer service, domain-specific assistants (e.g., AI volunteer assistant), and real-time personalized recommendation. When the access subject shifts from humans to agents, infrastructure faces three core challenges: traditional observability cannot capture dynamic decision paths; multi-turn, long-horizon tasks demand efficient memory storage, retrieval, and recall; and automated high-frequency agent access threatens database stability and requires new access protection.

Agent Observability: From System Monitoring to Decision Process Understanding

Traditional observability targets deterministic systems with fixed call topologies, pre-instrumented metrics (QPS, latency, error rate), and clear success signals like HTTP 200. Agent execution is a dynamic decision process with four structural blind spots: runtime-generated execution paths defy pre-instrumentation; token growth with context complicates cost attribution; technical success ≠ semantic correctness; multi-turn interactions explode trace and context data volume. Therefore, agent observability must "understand" — answering which path was taken, whether the result is correct, where costs incurred, and whether underlying storage can sustain load.

The team adopted Langfuse + ClickHouse. Agents report traces, spans, generations, tokens, costs, and scores via Langfuse SDK. Langfuse provides trace tracking, evaluation (LLM-as-a-Judge + human calibration), prompt/dataset management, and cost analysis, with ClickHouse handling high-throughput writes and multi-dimensional analysis. The observability platform visualizes full call chains, step-by-step inputs/outputs, tool calls, tokens, and latency, making dynamic processes replayable and evaluable.

Agent observability data exhibits "massive writes × semi-structured × multi-dimensional analysis". Traces contain large natural language text and JSON requiring long-term retention; metadata and tool calls evolve rapidly; queries naturally involve time windows, multi-dimensional filters, and metric aggregations. ClickHouse's columnar storage, sparse indexes, multi-level compression, and tiered storage reduce storage cost to 1/10 of traditional solutions . Native JSON/MAP types balance flexible ingestion and multi-dimensional queries. Query-side IO pruning, execution engine, and sparse/primary indexes support time-window filtering, joins, and high-throughput aggregations.

Bottlenecks in Community Langfuse at Scale

Single-shard ClickHouse : storage, writes, and large queries limited by single-node capacity.

Ingestion path : reliance on S3 historical replay and RMW (read-modify-write) merges causes read amplification, massive ClickHouse point queries, Redis ingestion queue buildup, and S3 upload latency.

Query path : Traces table primary key at day granularity forces full-day scans for ID lookups; ILIKE on input/output requires full-table scans; multi-table joins (traces, observations, scores); weak point-query performance for multimodal logs in columnar format.

Distributed Architecture Overhaul

Extended single-shard to multi-shard ClickHouse cluster, physically decoupling storage, writes, and query analysis. Combined with multi-replica HA and cold-data tiering to support massive agent observability data.

High-Throughput Ingestion Pipeline

Added Proxy module compatible with Langfuse V4 OTel interfaces. Data flows: Proxy → Pulsar → Sinker → ClickHouse. This yields 10× throughput improvement over native ingestion.

Query-Side Optimization with Langfuse V4

Migrated from trace-centric dual-entity model to observation-centric single-entity immutable model. SDK propagates trace common attributes to each observation, enabling wide-table queries on observations, reducing joins and updates. Added row-level indexes and late-filtering in the kernel to boost multimodal point-query performance and cut unnecessary CPU consumption.

Standardized Reporting for Downstream ML

Trace data feeds SFT and model distillation. To combat inconsistent reporting formats across Langfuse SDK versions, the team defined standardized reporting specs, a dedicated SDK, and an "reporting skill" that lets AI generate compliant instrumentation code.

Agent Memory Base: From OLAP Database to Remote Memory Service

Memory transforms historical information into context for current decisions. On user request, the system recalls task state, historical events, and user preferences from short-term and long-term memory, assembles them into the prompt, and passes to model reasoning and tool calling. This enables cross-turn continuity, personalization, historical result reuse, and continuous experience/feedback accumulation.

Four-Layer Memory Architecture

Short-term memory : current dialogue, task state, intermediate results.

Skill : task-type steps, rules, workflows.

Mem0 : decides what to remember, update, recall — user facts, preferences, experiences.

Memory Storage : long-term persistence.

Algorithms focus on extraction, filtering, updating; infrastructure focuses on storage and serving.

Remote Storage Necessity

Agents run in sandboxes where instances may be destroyed, restarted, scaled, or migrated — memory must survive instance lifecycle. Multiple agents often share documents, rules, code, or RAG corpora, requiring shared remote access.

Unified Retrieval in OLAP

Extended OLAP with BM25 full-text search, ANN vector search, and hybrid search, leveraging native aggregation and ranking to complete complex filtering, retrieval, statistics, and sorting in a single SQL.

DiskANN for Billion-Scale Vectors

Adopted disk-based DiskANN: PQ quantization compresses vectors; compressed vectors in memory, graph and full-precision vectors on SSD. Compared to HNSW, memory usage drops >90% while QPS decreases ~40% .

High-QPS Routing for Multi-Agent Online Serving

Built on WeChat's backend framework: hash routing by agent ID isolates workloads across nodes; supports fine-grained rate limiting, horizontal scaling, and read-write replica separation to meet high QPS, low latency, and stability.

Two Service Forms

Memory API/SDK : generic update/delete/access interfaces backed by OLAP vector database.

File-system memory service : FUSE wraps database operations as filesystem access, mounted into agent sandboxes. Agents read/write context and memory like files, avoiding new API learning curves.

The file-system view uses OverlayFS with User, Group, Agent three-layer directory trees. Different (User, Group, Agent) combinations yield distinct directory views merged via upper-layer-overrides-lower-layer into a unified sandbox tree. Shared memory placed at Agent layer; personalized memory at User layer. Q&A revealed a base/delta layering: delta uses local SQLite cache; reads check local first, fallback to remote; writes go local then async push to remote.

Agent Access Protection: From Human Access to Agent Access

Human database access is limited, intentional, accountable, and slow. Agent access can be unlimited, rule-driven, automatic, hard-to-attribute, and a tiny error can amplify instantly. Risks include over-privilege, SQL injection/malicious payloads, audit accountability gaps, and covert data leaks.

Core principle: treat agents as humans. Every database-accessing agent registers upstream, obtains DB permissions, and carries identity + dynamic tokens for auth and audit logging. Identity-based default rate limiting, rapid ban/intercept on anomalies.

Q&A Highlights

Q1: Langfuse single-shard → multi-shard?

Full architecture upgrade based on Langfuse V4. Ingestion path hashes directly into multi-shard ClickHouse. Adapted Langfuse Web queries to access distributed tables, with per-project sharding and permission control.

Q2: File-system memory storage implementation?

OLAP as remote storage; FUSE wraps as AgentFS mounted into sandbox. File reads translate to DB queries. Base/delta layering with local SQLite cache in delta: read local first, miss → fetch remote; write local first, then active push to remote.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

ObservabilityAccess ControlClickHouseVector SearchOLAPAgentic AIFUSEMemory SystemsLangfuseDiskANN
DataFunTalk
Written by

DataFunTalk

Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.