Scaling Enterprise AI Agents: Self‑GC Context Governance and NEX Sandbox
The AICon 2026 talks detail how Xiaohongshu’s engineering team tackles long‑context agent challenges by introducing Self‑GC’s ContextObject model, side‑channel commit mechanisms, a NEX sandbox for runtime isolation, and an L0‑L2 memory hierarchy, achieving 10‑20% token savings with over 90% impact‑free execution.
Self‑GC: A Prefix‑Cache‑Constrained Multi‑Turn Agent Context Governance Solution
As AI coding expands into broader enterprise collaboration, the bottleneck shifts from model instruction understanding to keeping long‑running agents stable in complex internal environments. Xiaohongshu’s quality‑efficiency R&D team presents Self‑GC, which models context as ContextObject and employs a Side‑channel Plan/Commit mechanism to make context objects indexable, recoverable, and governable throughout their lifecycle, thereby reducing token costs and mitigating long‑term “brain‑fog” degradation.
Cost and behavior degradation: Quantifies concurrent capacity pressure for a thousand‑person enterprise, explains “repeat‑loop” and tool‑call dead‑loop phenomena, and introduces Avg Input Tokens as the core governance metric.
Three engineering strategies: Compares lowering growth slope, In‑Run Pruning, and aggressive Compact, and establishes an Active View‑centric lossless optimization path.
Self‑GC runtime architecture: Details the identity and boundary definition of ContextObject, and the three‑phase commit flow—Async Plan, Rehearsal, Delayed Commit—that balances pruning efficiency with cache‑economy accounting.
Data flywheel and deployment insights: Shows how Self‑GC acts as an online session‑cleaning mechanism that feeds Benchmark/SFT/RL data, delivering a 10%‑20% reduction in input tokens while maintaining a 90%+ no‑impact rate in production.
Seal: Enterprise‑Scale AI Personal Assistant from Zero to Full‑Staff Coverage
The second talk explains the challenges of deploying an AI personal assistant across a large organization: controlling high token costs, preventing “forgetting” in long‑term tasks, and ensuring security and data isolation. The solution combines the NEX enterprise sandbox for runtime isolation, Self‑GC with Auto routing to halve costs, and an L0‑L2 layered memory system that transforms raw interaction logs into a searchable wiki.
AI Native project operation: Reviews the rapid rollout path of 3‑day internal testing, 1‑week public testing, and 2‑week full deployment, highlighting decision logic from MVP validation to scaling.
Enterprise‑grade security and cost dual‑solution: Describes NEX sandbox’s cluster and data isolation, and how Self‑GC’s context object governance together with SealRouter’s automatic model routing achieve cost reduction by 50%.
Memory enhancement and knowledge loop: Analyzes the L0 (short‑term context) to L2 (LLM‑backed wiki) hierarchy, solving audit‑hard “flush” and “dreaming” issues, and enabling traceable, filterable, and continuable memory.
AI ecosystem and next‑generation thinking: Presents Skill Hub and Cowork Studio as platforms for skill sharing and artifact publishing, and envisions a shift from passive response agents to goal‑driven, proactive, planning and collaborative intelligent agents.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Xiaohongshu Tech REDtech
Official account of the Xiaohongshu tech team, sharing tech innovations and problem insights, advancing together.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
