EvoOntology: Self-Evolving Ontology Boosts Data Agent Accuracy 17.8% While Cutting Tokens 20%

Renmin University's EvoOntology introduces a self-evolving ontology layer for data agents that transforms static semantic layers into runtime services, using execution traces to continuously patch content, tool, and schema layers via a builder agent and evolution agent with validation gates, improving trajectory-wise accuracy by 17.8% and reducing total token usage by 20% across multiple benchmarks.

DataFunTalk
DataFunTalk
DataFunTalk
EvoOntology: Self-Evolving Ontology Boosts Data Agent Accuracy 17.8% While Cutting Tokens 20%

Problem: Static Semantic Layers Hurt Data Agent Performance

Enterprise data agents need to understand business concepts like "Revenue" or "Customer" across disparate systems. The common approach is to pre-load a large semantic layer (table schemas, field descriptions, metrics, entity relationships) into the model's context. However, as data scale grows, this context becomes bloated with irrelevant information. Experiments on DDR-Bench with Claude-Sonnet-5 show that injecting a static semantic layer reduces Trajectory-Wise accuracy from 72.5% to 57.5%. Similar degradation appears on GPT-5.6-sol (68.5% → 65.5%). The core issue: agents waste tokens re-filtering a massive, static knowledge dump for each task.

Solution: Ontology as a Runtime Queryable Service (MCP Server)

EvoOntology restructures the ontology into three layers and serves it via an MCP (Model Context Protocol) server so agents query only what they need:

Content Layer : Stores business semantics — Terms (e.g., Revenue, Customer), Mappings (concept → physical tables/fields), Constraints (usage rules), and Evidence (supporting data samples).

Schema Layer : Manages the ontology's own structure, allowing new relationship types to be added.

Tool Layer : Defines how agents browse and resolve concepts (browse, resolve APIs). Agents receive a lightweight Manifest at task start, then fetch relevant pieces on demand.

This shifts ontology from a prompt prefix to an independent runtime service within the agent harness.

Self-Evolution Loop: From Execution Traces to Validated Patches

Initial Construction by Builder Agent

A Builder Agent ingests a batch of real tasks, extracts recurring entities, metrics, operations, and filters, then probes the underlying data (types, values, joins) to propose initial Terms, Mappings, Constraints, and Evidence. Each entry retains its Evidence for traceability.

Continuous Evolution by Evolution Agent

As agents run tasks, they leave execution trajectories (successes and failures). The Evolution Agent analyzes these traces for recurring failure patterns and attributes each to a layer:

Wrong field mapping → Content Layer patch

Correct knowledge exists but unretrieved → Tool Layer patch

New business relationship lacks representation → Schema Layer patch

The system generates a localized Patch (Candidate Ontology) rather than rebuilding the whole ontology.

Validation Gate (Parent vs. Candidate)

Before acceptance, the Candidate and Parent ontologies are evaluated on the same verification task set with identical model, decoding parameters, and budget. Only if the Candidate outperforms the Parent is the patch applied. Ablation confirms this gate is critical: removing it drops DDR-Bench Trajectory-Wise from 89.5% to 78.3%; removing attribution (layer diagnosis) also degrades results.

Experimental Results

Accuracy Gains Across Six Models

On DDR-Bench, InsightBench, and BIRD, EvoOntology improves Trajectory-Wise accuracy by an average of 17.8 percentage points over the no-ontology baseline. Notable per-model results:

Claude-Sonnet-5: 72.5% → 81.3%

GPT-5.6-sol: 68.5% → 93.5%

Four-model suite (GPT-5.5, GPT-5.6-sol, Claude-Sonnet-5, Claude-Opus-4.8): baseline 69.5% → evolved 89.5%

Token Efficiency: Total Tokens Drop ~20%

Although per-turn input tokens rise (3.2K → 4.6K), the number of turns per task falls sharply (14.6 → 8.4), yielding a total token reduction from 52.6K to 42K (~20%) . The ontology amortizes verified knowledge, eliminating repeated exploration (table search, field inspection, join trials).

vs. Episodic Memory

Adding a memory that retrieves past trajectories lifts accuracy to 75.8%; EvoOntology reaches 89.5%. The difference: memory stores task-specific history; ontology distills experience into queryable, combinable structure — closer to durable business rules.

Model-Specific Ontology Divergence

Starting from the same initial ontology, four models evolved independently. Pairwise Jaccard similarity of accepted Terms never exceeded 0.62 . Cross-model transfer (using Model A's evolved ontology on Model B) degrades performance by at least 6.6 points, averaging 10.9 points. This suggests ontology organization (what to expose, tool granularity) interacts with model behavior.

Layer Contribution Breakdown

Accepted evolutionary improvements attributed: Tool Layer 57% , Content Layer 34%, Schema Layer 9%. Over half the gain comes from how agents use the ontology (query timing, granularity, tool design), not just adding more knowledge.

Limitations & Real-World Gaps

Experiments run on DDR-Bench, InsightBench, BIRD. Real enterprises add permission control, conflicting business rules, multi-user edits, audit trails, and long-term version governance. EvoOntology's Parent/Candidate validation and rollback mechanism acknowledges that "self-evolving" is not fully autonomous — it automates a verified maintenance loop.

References

Paper: https://arxiv.org/abs/2609.15779 (EvoOntology: A Self-Evolving Ontology Layer for Data Agents)

Code: https://github.com/ruc-datalab/EvoOntology (RUC-DataLab / EvoOntology)

EvoOntology reframes ontology from a static document into a runtime component that agents both consume and maintain — closing the loop between execution experience and semantic knowledge.
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Semantic LayerMCP ServerData AgentsSelf-Evolving SystemsAgent TrajectoriesDDR-BenchEvoOntologyOntology Evolution
DataFunTalk
Written by

DataFunTalk

Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.