Ontology in Production: 14 Chinese Enterprises Share Hard-Won Lessons
DACon 2026 Beijing showcases 14 presentations from companies like Ant Group, Li Auto, and Shopee detailing how they successfully deployed ontology in production systems, covering architectural choices, semantic layer engineering, knowledge graph integration, and measurable business outcomes across finance, automotive, and e-commerce.
Why Ontology Deployment Remains Rare in China
The article opens with five core reasons why ontology (本体) has been discussed for over a decade in Chinese tech circles but few enterprises have successfully run it in production:
No unified standard: Palantir, Snowflake, Databricks, Apache Ossie, and Agent Memory each take different approaches — differences lie not in whether they use graphs, but in how knowledge is represented, executed, produced, and validated.
Tendency to become "big and comprehensive": Enterprise-wide ontologies built in one go become obsolete fastest; they bloat, break, and lose value when business changes.
Irreversible implementation path: Choosing between a semantic mapping layer (fast to launch, poor multi-hop performance) and a materialized object store (fast queries, but faces sync, consistency, cost, and "source of truth" challenges) is largely a one-way decision.
Definition conflicts: Business lines define metrics like "active user" differently; forced unification distorts understanding, while lack of unification leaves agents confused.
Depreciation after delivery: Without continuous iteration mechanisms, knowledge assets start depreciating from day one.
Section 1: Enterprises with Ontology Already Running in Production (7 Talks)
01 Ant Group: Financial Data Knowledge Ontology — From "Querying Data" to "Understanding Business"
Speaker: Cui Weimin, Senior Data Technology Expert, Ant Group.
Core problem: Business lines defined "active user" differently. Ant's solution: retain differences in the ontology, let agents fetch on demand.
Architecture:
LangGraph 6-node state machine: InputParser → Comprehender → EntityResolver → Reconciler → Synthesizer → Verifier — automates ontology synthesis and validation.
SQL Understander (sqlglot + regex): Reverse-extracts field semantics and cross-table relationships from historical task code, creating a "code-as-annotation" auto-generation pipeline.
Reconciler: 7-level cost-progressive fusion for cross-table same-name fields (auto-merge → human confirmation), core principle: flag conflicts but don't force merge.
Verifier: Pydantic + business rules + deep-pit detection; failures trigger L4 retry or downgrade warning.
Service layer — mutual-index architecture: Synchronously retrieves three knowledge types: KG_cs (business-view rules, e.g., Huabei's "active user" definition), KG_fr (entity-relation graph, e.g., "User A — holds — Huabei limit 50k"), RC (raw document evidence). Results fused, deduplicated, packaged as KnowledgePackage.
Results: Cross-opportunity identification reduced from 3–5 days to hour-level; anomaly attribution accuracy >85%; supports multiple marketing strategies in production.
"Ontology is not a dictionary; it's the textbook that lets the system 'understand business'."
02 Li Auto: Ontology-Driven Agent — From Query & Analysis to Sales Decisions
Speaker: Liang Wei, Senior Algorithm Engineer, Li Auto.
Background: Post-training achieved near 100% NL2SQL accuracy, but faced long knowledge-prep/training cycles, difficulty extending to new domains, and brittleness when business definitions changed.
Approach: Re-organized enterprise data with ontology so objects, relations, metrics, and business rules become semantic structures agents can understand and query.
Evolution path: NL2SQL → MCP → Skill → built-in Agent — four iterations with trade-offs. Sales decision support built on three layers: business fact layer, sales cognition layer, scenario execution layer.
03 Li Auto: Semantic Layer as Context — Production-Grade Agent Semantic Engineering
Speaker: Qian Han, Big Data Platform Lead, Li Auto.
Thesis: Production agent bottleneck is not model capability but context quality — whether agents get fragmented info scattered in prompts/tools/logs or semantically governed, traceable, structured context.
Case: Spark/Flink fault diagnosis. Knowledge split into four layers: Ontology, OSI, Playbook, General Execution Layer. Agents cross-reference logs, metrics, data warehouse to produce evidence-backed root-cause judgments; diagnostic gaps become versioned, verifiable, regressible semantic assets.
04 Tianyi Payment: Ontology Deployment Fork — Semantic Mapping vs. Object Storage
Speaker: Xu Dehua, Risk Management Director, Tianyi Payment.
Core dilemma: Is ontology a "translation layer" or a "data store"?
Mapping only: Metadata layer, translates to warehouse SQL at query time — no data movement, fast launch, no redundancy; but multi-hop traversal fails, definitions bound to underlying table structures, no place to land object state/execution results.
Materialized object store: Instances, attributes, relationships physically stored — fast queries, supports state & decision closure; but requires real-time sync, incremental consistency, storage cost, and "who is the truth" resolution.
Decision: Basically irreversible. Upstream issue: concept definitions and real business data shouldn't be mixed in one layer.
Solution: Hybrid — core objects materialized, long-tail kept as mapping; boundary fixed in design, not ad-hoc. Four selection criteria: query pattern, latency needs, need for state/write-back, change frequency.
Agent-facing data supply: Minimum necessary context, governed access interfaces, fallback conditions, risk quantification.
Cross-domain reuse & acceptance: Five-piece delivery per business domain: ontology fragment, rule library, case library, verified queries, evaluation set.
05 Shopee: Graph-Platform-Based Anti-Fraud Ontology Modeling & Intelligent Decisioning
Speaker: Zhang Songqing, Graph Platform Lead, Shopee Data Infra.
Problem: Traditional tabular models struggle with coordinated disguise across accounts, devices, transactions in e-commerce promotions and credit.
Platform: One-stop graph platform — graph DB + graph compute + graph learning synergy. Objects, relations, risk rules modeled as ontology → dynamic anti-fraud knowledge network.
Two real scenarios:
Monee Credit user relationship graph & credit risk control: Pre-/in-loan identification uplift; graph DB enables real-time eligibility approval for online loan applications.
Shopee order promotion knowledge graph & fraud detection: Graph compute + RGAT (Relational Graph Attention Network) identifies black-industry syndicates and anomalous transaction chains.
Engineering challenges: Large-scale heterogeneous relations, real-time updates, feature serving, training efficiency; solved with feature pre-fetch and batch-processing optimization.
06 Runhe Software: Domain Knowledge Engineering with Ontology — Financial Testing Case
Speaker: Yang Dongrui, AI Architect, Runhe Software.
Pain points: LLM direct generation suffers: methods not reusable, results not traceable, quality not evaluable.
Solution: Test Ontology + Test IR (Intermediate Representation). Business ontology precipitates testing methodology; Test IR connects requirements, outlines, cases. Agent handles semantic understanding & method selection; tools handle data calc, case assembly, validation — forming traceable chain: requirement parsing → IR → method application → case generation.
Four-route comparison: Direct LLM (fast but black-box), fixed workflow/prompt (single-scene but tight coupling), RAG knowledge retrieval (supplements knowledge but can't express method applicability constraints & output constraints), Ontology + Test IR (structured, suitable for long-term governance).
07 360: From "Querying Data" to "Understanding Business" — Progressive Semantic Layer Evolution & Quantitative Evaluation
Speaker: Guo Chaoyang, Lakehouse Data Agent Evolution Tech Lead, 360.
Gap: LLM writes SQL at 85% on academic benchmarks; in real warehouse (hundreds of tables, 10k+ fields, tribal definitions) accuracy collapses, and errors are opaque.
Strategy: "Half-finished cold start, grow while running"
Dual-path hybrid routing: Semantically complete → deterministic compilation engine (AST main path); only tables/docs → RAG fallback. Every result carries confidence score: ≥0.95 auto-enters reports; ≤0.65 system prompts "low confidence, suggest human review".
Thin API → Thick Skill → Ontology metadata: High-frequency usage patterns reverse-precipitate as semantic assets; Skills eventually degrade to pure orchestrators, knowledge fully provided by ontology layer.
Three-level evaluation + feedback flywheel: L1 programmatic result-set comparison → L2 LLM Judge semantic judgment → L3 human fallback. User thumbs-up (after anti-pollution filtering) auto-sampled into regression test set.
Current scope: Internal big-data cluster ops, S3/PoleFS storage ops metric analysis & attribution; Ops data-center ops, document search, membership middle-platform across departments.
Section 2: Ontology & Semantic Layer — Engineering Methodology & Selection Map (5 Talks)
08 Yuepoint Tech: Next Stop for Semantic Layer Is Business Ontology — 12 Years, ~100 Enterprises, Gains & Losses
Speaker: Ren Xinqi, Founder/CEO, Yuepoint Tech.
Key distinction: Domain ontology ≠ another name for warehouse semantic layer; it's an independent enterprise business world model — organizing business objects, relations, rules, executable actions.
Pitfalls observed: Terminology & metric definition inconsistency; "big and comprehensive" ontologies obsolete first; business/algorithm/data tri-party cognitive misalignment.
Methodology: Scenario-first lightweight ontology modeling; LLM-assisted knowledge extraction + expert review closed-loop iteration.
Ontology constraining LLMs: Concept verification & generation constraints across RAG & Agent full chain — unify definitions, reduce hallucination.
Case: Manufacturing quality traceability — cross CRM/MES/ERP/PLM trace from multi-person days to one person 5 minutes; yield improved 5%.
09 Datastrato: Open Semantics, Ontology & Metadata — Three Pillars of Agent Context
Speaker: Du Junping, Founder/CEO, Datastrato; Apache Gravitino initiator.
Three-layer progression:
Unified metadata: Apache Gravitino cross-cloud, cross-engine, multi-modal unified metadata management — consistent description & access interface for data assets.
Open semantic layer: Standardize metrics, dimensions, business definitions — eliminate "one metric, multiple definitions".
Ontology: Model business entities, relations, rules — enable agents to precisely map natural-language business intent to real data.
Two context-building strategies compared: Top-down (task-driven, fast but creates context islands) vs. Bottom-up (metadata-governed, solid foundation but can't infer business meaning). Sustainable hybrid: business-scenario-driven modeling + unified metadata connecting business concepts, data assets, governance rules + continuous iteration via business validation & agent feedback.
10 Cloud器 (CloudQi): No Unified Ontology Standard — Five Routes, Which to Follow?
Speaker: Guan Tao, Co-founder & CTO, Cloud器.
Five mainstream routes: Palantir, Snowflake, Databricks, Apache Ossie, Agent Memory. Real difference not "graph or not" but how knowledge is represented, executed, produced, validated.
From DIKW: Deconstruct knowledge & ontology evolution; compare five routes; give enterprise path from data engineering to knowledge engineering.
11 Dongchedi (懂车帝): AI-Native Data Semantic Platform — Ontology as Strong-Constraint Knowledge Layer
Speaker: Miao Zhiyong, Data Warehouse Lead, Dongchedi.
Problem: LLM direct answers on business data → hallucination, instability, low confidence. Text2SQL relying solely on model training or text recall can't guarantee accuracy/consistency.
Semantic layer value: Strong-constraint knowledge between LLM and data assets. Structurally manages metrics, dimensions, entities, business processes, physical implementations, SQL rules — lets AI understand business, generate SQL, explain results within explicit boundaries. vs. single text KB, RAG vector recall, full metadata recall: semantic layer has clearer query constraints, explainability, lower inference cost.
Platformization solves: Stale definitions, physical table updates, local version inconsistency, scattered knowledge maintenance — platform as single source of truth, continuously manages semantic definitions, versions, mappings, consumption governance.
12 Yiche (易车): From DataAgent to Action Closure — Ontology as One of Three Foundational Layers
Speaker: Wang Linhong, Data Platform Lead, Yiche.
Evolution: ChatBI → DataAgent → DataWork. Three-layer foundation covering trusted data, semantics, ontology & knowledge, MCP & Harness. Query evolution + proactive follow-through connect analysis, action, result verification — forming enterprise data intelligence progression from trusted querying → assisted analysis → controlled execution.
Section 3: Frontier Exploration & Cross-Industry Extension (2 Talks)
13 Renmin University: Self-Evolving Data Agents — Data–Ontology–Agent Co-Evolution
Speaker: Zhang Shaolei, Assistant Professor, RUC School of Information.
Problem: Existing data agents face alignment gaps among business intent, data semantics, tool execution on heterogeneous data (tables, databases, docs). Fixed workflows, RAG, static KGs lack continuous adaptation to open tasks & dynamic data.
Framework: Ontology as executable semantic middle layer — unifies domain concepts, tool capabilities, data schemas. Based on agent interaction trajectories: failure attribution → knowledge correction → pairwise evaluation gating.
Key insight: Typical agent errors stem not from model hallucination but from data object identification, field semantic mapping, tool selection errors.
Three-layer ontology: Content Layer (business concepts/metric definitions/entity relations), Tool Layer (tool capabilities & constraints), Schema Layer (logical concept → physical table/field mapping).
Evolution loop: Interaction trajectory failure attribution → candidate modification generation → hold-out validation pairwise evaluation gating → versioned release / reject rollback. Forms continuous loop: "data exposes problem → agent produces trajectory → ontology precipitates knowledge → agent re-validates".
Status: Running 6 months; 2000+ registered users, ~400 deep users; 300+ universities/institutes requested trial.
14 Global Luxury Retailer × Shinx (数新智能): Apache Ossie Ontology Spec & Semantic Enhancement
Speakers: Tan Zong (APAC Big Data Lead, global top luxury retailer) × Yuan Panfeng (CTO, Shinx).
Context: LLMs evolving from single-point Copilot to Agent paradigm; enterprise data scattered across multi-cloud multi-brand, strict compliance, fragmented AI apps — agents "no data available".
Two-part talk: Business perspective (pain points: data silos, compliance walls, fragmented AI limits, Agent-Ready data foundation demands) + Technical perspective (multi-cloud AI-native platform DataAgent architecture, enterprise semantic layer, sandbox security, closed-loop iteration).
Results: End-to-end data delivery efficiency +70%; Agent task diagnosis accuracy +50%; via semantic knowledge base, code generation accuracy from 30% to 90%.
Engineering highlights: Apache Ossie ontology spec, CodeWiki semantic enhancement to补齐 business knowledge gaps, multi-version compatibility/fault-tolerance.
Closing Synthesis
The article concludes that ontology deployment in China is scarce because it demands a chain of upfront judgments: separate concept vs. instance layers? mapping vs. object store? force-unify definitions or preserve differences? one-off delivery or evolving infrastructure? how to prove superiority over prior approach? No standard answers — only trade-offs made under each enterprise's constraints. The value lies in those trade-offs and their costs. The 14 talks at DACon 2026 Beijing (Oct 23–24, Hilton Beijing) represent the densest public sharing to date of Chinese enterprises' ontology production experiences.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunTalk
Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
