EvoOntology: Self-Evolving Ontology Layer Lets Data Agents Maintain Their Own Semantic Knowledge

EvoOntology introduces a self-evolving ontology layer for data agents that transforms static semantic layers into a runtime service, using a three-layer architecture and an evolution loop where agents' execution traces drive continuous ontology updates via builder and evolution agents with validation gates, reducing token usage and improving accuracy across benchmarks.

DataFunSummit
DataFunSummit
DataFunSummit
EvoOntology: Self-Evolving Ontology Layer Lets Data Agents Maintain Their Own Semantic Knowledge

The article analyzes EvoOntology , a research system from Renmin University (RUC) that addresses the core maintenance problem of ontologies in enterprise Data Agent deployments. Traditional semantic layers are static, manually curated, and become stale as business definitions, schemas, and agent behaviors change. EvoOntology reframes the ontology as a runtime service that agents query on demand and continuously update from their own execution traces.

Problem: Static Semantic Layers Hurt Agent Performance

Experiments on DDR-Bench show that stuffing a full semantic layer into the prompt degrades accuracy. For Claude-Sonnet-5, baseline Trajectory-Wise accuracy is 72.5%; adding a static semantic layer drops it to 57.5%. Similar regressions appear on GPT-5.6-sol (68.5% → 65.5%). The reason: large context windows force agents to filter irrelevant knowledge each turn, increasing exploration steps and token cost.

Architecture: Three-Layer Ontology as an MCP Server

EvoOntology splits the ontology into three layers, exposed via an MCP (Model Context Protocol) server so agents browse only what they need:

Content Layer – business concepts (Terms), mappings to physical tables/fields, constraints, and supporting Evidence (e.g., sample values that justify a mapping).

Schema Layer – governs the ontology's own structure, allowing new relationship types to be added.

Tool Layer – defines how agents query the ontology (browse, resolve, etc.). Agents receive a lightweight Manifest at task start, then fetch relevant pieces on demand.

This turns the ontology from a prompt prefix into an independent runtime component inside the agent harness.

Bootstrapping: Builder Agent

An initial ontology is not purely hand-crafted. A Builder Agent ingests a batch of real tasks, extracts recurring entities, metrics, operations, and filters, then probes the underlying data (types, values, joins) to propose Terms, Mappings, and Evidence. Only verified mappings enter the initial ontology.

Continuous Evolution: Evolution Agent + Validation Gate

As agents run tasks, they leave execution trajectories . An Evolution Agent analyzes these traces for repeated failure patterns (wrong field mapping, missed retrieval, missing structural concepts). It attributes each pattern to a specific layer (Content, Tool, or Schema) and generates a local Patch rather than rebuilding the whole ontology.

Crucially, every Candidate Ontology is evaluated against a held-out validation set under identical model, decoding, and budget settings . Only if the new version outperforms the Parent Ontology is it accepted. Ablation confirms this gate matters: removing it drops DDR-Bench Trajectory-Wise from 89.5% to 78.3%; removing attribution also degrades results.

Experimental Results

Accuracy gains : Across six models, EvoOntology improves Trajectory-Wise by 17.8 percentage points on average vs. no-ontology baseline. Final evolved scores reach 89.5% (DDR-Bench), 93.5% (GPT-5.6-sol).

Token efficiency : On DDR-Bench, baseline agents use ~52.6K total tokens/task (14.6 turns). After evolution, total tokens fall to ~42K (8.4 turns) – a ~20% reduction – despite higher per-turn input (3.2K → 4.6K). The ontology eliminates repeated schema exploration.

vs. Memory : A ReAct baseline with retrievable memory reaches 75.8%; EvoOntology reaches 89.5%. Memory stores past trajectories; ontology distills them into queryable, combinable structure.

Model-Specific Ontologies

Four models (GPT-5.5, GPT-5.6-sol, Claude-Sonnet-5, Claude-Opus-4.8) evolved from the same initial ontology. Final Term sets show pairwise Jaccard similarity ≤ 0.62 . Cross-model transfer degrades performance by 6.6–10.9 points, indicating that optimal ontology organization (what to expose, tool granularity) is model-dependent.

Layer Contribution Breakdown

Accepted patches: Tool Layer 57% , Content Layer 34%, Schema Layer 9%. Over half the gain comes from improving how agents use the ontology , not just adding more facts.

Limitations & Outlook

Experiments run on DDR-Bench, InsightBench, and BIRD. Real enterprises add permission control, rule conflicts, multi-user editing, audit, and long-term version governance. The Parent/Candidate + validation + rollback loop shows that "self-evolving" is a controlled automation cycle , not unrestricted agent writes. The ontology becomes a living runtime component: agents use it to understand data, and their execution results continuously refine it.

Reference: EvoOntology: A Self-Evolving Ontology Layer for Data Agents — arXiv:2609.15779 (https://arxiv.org/abs/2609.15779) Code: RUC-DataLab/EvoOntology (https://github.com/ruc-datalab/EvoOntology)
EvoOntology overall architecture: Data Agent interacts with three-layer ontology via MCP tools, and evolution loop diagnoses failures from trajectories to generate patches.
EvoOntology overall architecture: Data Agent interacts with three-layer ontology via MCP tools, and evolution loop diagnoses failures from trajectories to generate patches.
Performance comparison: Baseline → Initial Ontology → Evolved Ontology on DDR-Bench, InsightBench, and BIRD.
Performance comparison: Baseline → Initial Ontology → Evolved Ontology on DDR-Bench, InsightBench, and BIRD.
Metric trends across evolution rounds for four models on three benchmarks.
Metric trends across evolution rounds for four models on three benchmarks.
Agent cost vs. performance: per-turn input tokens rise, but total turns and total tokens drop as ontology evolves.
Agent cost vs. performance: per-turn input tokens rise, but total turns and total tokens drop as ontology evolves.
Jaccard similarity of Terms between model-specific evolved ontologies, all below 0.62.
Jaccard similarity of Terms between model-specific evolved ontologies, all below 0.62.
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Semantic LayerLLM AgentsKnowledge GraphsontologyBenchmark EvaluationData AgentsSelf-Evolving SystemsEvoOntology
DataFunSummit
Written by

DataFunSummit

Official account of the DataFun community, dedicated to sharing big data and AI industry summit news and speaker talks, with regular downloadable resource packs.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.