EvoOntology: Self-Evolving Ontology Layer Lets Data Agents Maintain Their Own Semantic Knowledge
EvoOntology introduces a self-evolving ontology layer for data agents that transforms static semantic layers into a runtime service, using a three-layer architecture and an evolution loop where agents' execution traces drive continuous ontology updates via builder and evolution agents with validation gates, reducing token usage and improving accuracy across benchmarks.
The article analyzes EvoOntology , a research system from Renmin University (RUC) that addresses the core maintenance problem of ontologies in enterprise Data Agent deployments. Traditional semantic layers are static, manually curated, and become stale as business definitions, schemas, and agent behaviors change. EvoOntology reframes the ontology as a runtime service that agents query on demand and continuously update from their own execution traces.
Problem: Static Semantic Layers Hurt Agent Performance
Experiments on DDR-Bench show that stuffing a full semantic layer into the prompt degrades accuracy. For Claude-Sonnet-5, baseline Trajectory-Wise accuracy is 72.5%; adding a static semantic layer drops it to 57.5%. Similar regressions appear on GPT-5.6-sol (68.5% → 65.5%). The reason: large context windows force agents to filter irrelevant knowledge each turn, increasing exploration steps and token cost.
Architecture: Three-Layer Ontology as an MCP Server
EvoOntology splits the ontology into three layers, exposed via an MCP (Model Context Protocol) server so agents browse only what they need:
Content Layer – business concepts (Terms), mappings to physical tables/fields, constraints, and supporting Evidence (e.g., sample values that justify a mapping).
Schema Layer – governs the ontology's own structure, allowing new relationship types to be added.
Tool Layer – defines how agents query the ontology (browse, resolve, etc.). Agents receive a lightweight Manifest at task start, then fetch relevant pieces on demand.
This turns the ontology from a prompt prefix into an independent runtime component inside the agent harness.
Bootstrapping: Builder Agent
An initial ontology is not purely hand-crafted. A Builder Agent ingests a batch of real tasks, extracts recurring entities, metrics, operations, and filters, then probes the underlying data (types, values, joins) to propose Terms, Mappings, and Evidence. Only verified mappings enter the initial ontology.
Continuous Evolution: Evolution Agent + Validation Gate
As agents run tasks, they leave execution trajectories . An Evolution Agent analyzes these traces for repeated failure patterns (wrong field mapping, missed retrieval, missing structural concepts). It attributes each pattern to a specific layer (Content, Tool, or Schema) and generates a local Patch rather than rebuilding the whole ontology.
Crucially, every Candidate Ontology is evaluated against a held-out validation set under identical model, decoding, and budget settings . Only if the new version outperforms the Parent Ontology is it accepted. Ablation confirms this gate matters: removing it drops DDR-Bench Trajectory-Wise from 89.5% to 78.3%; removing attribution also degrades results.
Experimental Results
Accuracy gains : Across six models, EvoOntology improves Trajectory-Wise by 17.8 percentage points on average vs. no-ontology baseline. Final evolved scores reach 89.5% (DDR-Bench), 93.5% (GPT-5.6-sol).
Token efficiency : On DDR-Bench, baseline agents use ~52.6K total tokens/task (14.6 turns). After evolution, total tokens fall to ~42K (8.4 turns) – a ~20% reduction – despite higher per-turn input (3.2K → 4.6K). The ontology eliminates repeated schema exploration.
vs. Memory : A ReAct baseline with retrievable memory reaches 75.8%; EvoOntology reaches 89.5%. Memory stores past trajectories; ontology distills them into queryable, combinable structure.
Model-Specific Ontologies
Four models (GPT-5.5, GPT-5.6-sol, Claude-Sonnet-5, Claude-Opus-4.8) evolved from the same initial ontology. Final Term sets show pairwise Jaccard similarity ≤ 0.62 . Cross-model transfer degrades performance by 6.6–10.9 points, indicating that optimal ontology organization (what to expose, tool granularity) is model-dependent.
Layer Contribution Breakdown
Accepted patches: Tool Layer 57% , Content Layer 34%, Schema Layer 9%. Over half the gain comes from improving how agents use the ontology , not just adding more facts.
Limitations & Outlook
Experiments run on DDR-Bench, InsightBench, and BIRD. Real enterprises add permission control, rule conflicts, multi-user editing, audit, and long-term version governance. The Parent/Candidate + validation + rollback loop shows that "self-evolving" is a controlled automation cycle , not unrestricted agent writes. The ontology becomes a living runtime component: agents use it to understand data, and their execution results continuously refine it.
Reference: EvoOntology: A Self-Evolving Ontology Layer for Data Agents — arXiv:2609.15779 (https://arxiv.org/abs/2609.15779) Code: RUC-DataLab/EvoOntology (https://github.com/ruc-datalab/EvoOntology)
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunSummit
Official account of the DataFun community, dedicated to sharing big data and AI industry summit news and speaker talks, with regular downloadable resource packs.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
