Industry Insights 42 min read

12 Chinese Tech Giants Share Production Semantic Layer Architectures for AI Data Agents

DACon 2026 brings together 12 leading Chinese companies including Ant Group, Gaode, Zhihu, and Li Auto to detail how they built semantic layers that bridge LLMs and enterprise data warehouses, revealing concrete architectures, accuracy benchmarks, and lessons learned from moving beyond demos to production-grade Data Agents.

DataFunSummit
DataFunSummit
DataFunSummit
12 Chinese Tech Giants Share Production Semantic Layer Architectures for AI Data Agents

Conference Overview: DACon 2026 Beijing

On October 23-24, 2026, DACon 2026 Beijing hosted 12 sessions from companies that have implemented semantic layers at scale. The central theme: large models achieve 85% SQL accuracy on academic benchmarks but plummet to ~50% in real warehouses with hundreds of tables, thousands of columns, and tribal knowledge. The solution is a semantic layer — a "strong constraint knowledge" layer between LLMs and data assets.

Four Universal Hurdles

Semantic maintenance: Definitions expire, physical tables change, local versions diverge, knowledge scattered.

Execution reliability: LLMs generate SQL that fails silently; errors are undetectable without deterministic compilation.

Permission inheritance: Metadata visibility ≠ data access; most systems only handle the former.

Effect quantification: Academic benchmarks compare SQL text, not query results; they ignore business definitions, multi-turn dialogue, and hallucination detection.

Company Case Studies

1. 懂车帝 — AI-Native Semantic Platform

Speaker: 苗治勇 (Data Warehouse Lead). Defines semantic layer as "strong constraint knowledge" structuring metrics, dimensions, entities, business processes, physical mappings, and SQL rules. Compared to alternatives: vs. pure text knowledge base — clearer query constraints, better explainability, lower inference cost; vs. RAG vector retrieval — defines rules not similarity; vs. full metadata dump — prevents AI guessing across thousands of tables. Platform solves definition expiration, table changes, version inconsistency, scattered maintenance via unified fact source.

2. 知乎 — Three-Stage Evolution: MetricFlow → Knowledge Base → Lightweight Knowledge Graph

Speaker: 汤晋瑄 (Data Platform Engineer). Stage 1: dbt MetricFlow + BM25 retrieval. SQL generated by MetricFlow engine — standardized, type-safe, accurate. Limitations: complex, high cost for special metrics, inflexible, cannot leverage LLM fully, high config maintenance. Stage 2: LLM auto-generates metric definitions stored as Markdown; closed loop: LLM generate → knowledge base ingest → enhanced retrieval → LLM use. Scale: 1,400+ metrics, 300+ wide/certified tables, 25,000+ SQL templates, business docs. Stage 3: Knowledge graph for multi-hop relations. AI-native, lightweight: AI-led construction, maintenance, querying; no heavy graph compute. Core tech: custom chunking (per metric, layered tables, per SQL statement, fallback), hybrid retrieval (full-text + vector + WRRF fusion + name matching), hierarchical matching (table → field), business priority (core wide tables > certified tables). Result: team penetration >65%, simple queries 100% satisfied, token consumption drastically reduced.

3. 蚂蚁集团 — Apache Ossie Semantic Layer: 100+ Models, 5,000+ Metrics

Speaker: 王小军 (Semantic Platform Lead). Problem: Data Agents write SQL but don't understand business. Solution: Open-source OSI protocol with "base standard + business extensions". Asset platform hosts single semantic source of truth (SSOT), unifying models, lineage, consumption — "define once, consume everywhere". Converts tables, BI models, queries into traceable human+machine co-built models. Forward dev pipeline + lineage monitoring for freshness. Subgraph retrieval, confidence scoring, ranking for agent consumption. Current: 100+ semantic models, 5,000+ metrics, multi-scenario consumption (human, Agent, platform). Challenges: OSI native standard lacks enterprise features (multi-grain metrics, equivalent metrics, cross-model refs) — solved by controlled extensions; model drift — solved by full-chain monitoring (feedback, patrol, change analysis) + forward dev pipeline embedding semantic iteration into data dev lifecycle; consumption accuracy at scale — solved by "hybrid retrieval + progressive disclosure" (keyword+vector+tags+graph, confidence scoring, re-ranking, return task-specific subgraphs).

4. 理想汽车 — Semantic Layer as Context for Production Agents

Speaker: 钱瀚 (Big Data Platform Lead). Thesis: Production Agent bottleneck is context quality, not model capability. Uses Spark/Flink fault diagnosis example. Splits scattered Prompt/tool knowledge into Ontology, OSI, Playbook, generic execution layer. Agent cross-references logs, metrics, warehouse for evidence-based root cause. Diagnoses gaps become versioned, verifiable, regressible semantic assets. Discusses lightweight modeling, query constraints, exploration vs. determinism trade-offs. Challenges: manual ontology/playbook effort limits cross-business replication; need standardized auto-generation. Future: permission control, execution tracing, evidence status, result validation, multi-environment adaptation, governable/observable/extensible Agent control plane.

5. Datastrato — Three Pillars of Agent Context: Unified Metadata, Open Semantic Layer, Ontology

Speaker: 堵俊平 (Founder/CEO, Apache Gravitino PMC). Layered design: Unified Metadata (Gravitino: cross-cloud, cross-engine, multimodal); Open Semantic Layer (standardized metrics, dimensions, business definitions — eliminate multiple calibers); Ontology (business entities, relationships, rules — map natural language intent to real data). Layers build progressively. Two build paths: top-down (app-driven) fast value but creates context silos; bottom-up (metadata-driven) broad coverage but cannot derive business meaning from technical structure alone. Sustainable hybrid: business scenarios drive modeling, unified metadata connects concepts/assets/governance, continuous iteration via business validation and agent feedback. Challenges: fragmented semantic modeling across apps; business meaning not auto-derivable from tech structure; lack of standards for multimodal/cross-cloud/cross-engine metadata; context freshness — solved by unified metadata base + open semantic layer + ontology rules/boundaries + scenario-driven modeling + feedback iteration.

6. 高德 (Talk A) — Semantic Governance, Attribution Diagnosis, Management Decision

Speaker: 钟雨洁 (AI Application Data Analyst). Local life business: 20+ industries, vastly different models/calibers/focus. Manual analysis slow, experience not reusable. Three-layer Agent architecture: Semantic Governance unifies business calibers/knowledge; Attribution Diagnosis locates metric changes and business root causes; Management Decision converts conclusions to actions. Knowledge architecture: Edit-Compile-Publish-Consume. Edit layer manages industry knowledge, metric rules, operational experience. Compile layer transforms to business entities, strong-typed relations, atomic knowledge statements. Publish layer generates immutable versions with content hashes. Consume layer serves unified interface to analysis/decision Agents. Query pipeline: capability catalog → identity resolution → relation retrieval → task routing → knowledge assembly; ambiguity, missing, expired, truncated explicitly returned. Results: ~66,000 traceable knowledge statements, 198 consumption paths (186 direct-service). Key division: engineering system guards numbers, identities, scopes, versions; model handles business interpretation, synthesis, natural language. Numbers and attribution computed by deterministic engine; model explains within fixed facts/evidence boundaries. Reports closed via structured contracts, execution receipts, content fingerprints, file hashes; Guard intercepts missing signals, unsourced numbers, causal leaps. Efficiency up 64x, business acceptance 100%.

7. 高德 (Talk B) — Real Architecture Migration: From "Answering Right" to "Seeing Through"

Speaker: 林宇航 (Data Product Expert). Pain: dashboards never finished, long-tail uncovered; traditional query only answers "how much", not "why/what now". Did not stack larger models; transformed multi-year BI base into routable, domain-knowledge-loadable Agent base. Real migration in two months: MVP 39 tables → 89.4% accuracy; full 331 tables → end-to-end accuracy crashed to 50%; switched from RAG+Workflow to Skill architecture → recovered to 95%. Scale challenge: single-domain MVP works, but full-scale exposes RAG+Workflow flaws — one-shot knowledge retrieval, multi-step info loss, no cross-table complex calc. Skill (progressive disclosure) with on-demand retrieval + AI self-reflection correction restored accuracy. Core judgment: semantic modeling is key, not model itself. Gartner L2→L3→L4 mapping: currently L2→L3 transition, L4 autonomous decision blank, 12-18 month window. Dual-engine: 小豪 ChatBI for L2普惠; Cici Skill+Agent for deep L3-L4. AI-friendly knowledge base: three sources (auto metadata, human-fed business knowledge, routing rules). Evaluation flywheel: AI-friendly standard score, NL2SQL EM+EX, end-to-end vs OpenAI Trace Grading.

8. Datastrato (Talk B) — Enterprise-Grade Trusted Query Engineering Loop

Speaker: 李明皇 (Senior Engineer, Apache Gravitino PMC). Four hurdles (same as above). Engineering loop: modeling → compilation → maintenance → evaluation → governance. Unified metric semantics compiled to executable SQL. Continuous semantic drift detection, evaluation-set-driven regression. Business domain + tag policies for fine-grained permissions. Uses Apache Gravitino with thousands of metrics. Demo: Agent doesn't guess calibers or free-generate SQL; calls controlled deterministic tools for asset discovery, semantic parsing, query compilation, execution. Returns result plus semantic basis, data sources, tool call chain — explainable, reproducible, auditable. Ambiguity/unsafe joins trigger clarification or refusal. Semantic Generator uses metadata, query history, domain knowledge to produce auditable proposals; combined with deterministic validation, human review, drift detection, regression testing. Closing line: "Insights trustworthy, then operational actions dare to automate."

9. 奇虎360 — Half-Baked Cold Start, Grow While Running

Speaker: 郭朝阳 (Lakehouse Data Agent Evolution Lead). Reality: business won't wait 6 months for perfect semantic layer. Most enterprises only have scattered table docs and tribal knowledge. Core shift: "model writes SQL" → "engine writes SQL"; uncertain → model, certain → engine. But new problem: semantic layer incomplete. Solution: "half-baked cold start, grow while running". Dual-path hybrid routing: semantic-complete → deterministic compilation engine (LLM only extracts intent); only tables/docs → RAG fallback (hybrid retrieval: embedding + keyword + foreign-key graph). Routing decision: PreMatch + confidence scoring. Every result carries confidence score — user knows "this is computed, that is guessed". Confidence enables production use in high-stakes decisions (e.g., product ops analyzing retention impact of feature change): >0.95 auto-include in reports; <0.65 system prompts "uncertain, verify manually". Skill system: thin API → thick Skill → ontology metadata. Aligns with Snowflake Cortex Agent layers; end state: Skill degrades to pure orchestrator, knowledge fully from ontology. Three-level evaluation: L1 programmatic result-set comparison → L2 LLM Judge semantic judgment → L3 human fallback. Data flywheel: user likes filtered, auto-sampled into regression test set; 9 metrics tracked continuously. Supports internal big data cluster ops, S3/PoleFS storage ops, cross-dept (Ops, doc search, membership).

10. 小米 — From BIRD Global 3rd to Production: Semantic Layer as Evolution Fuel

Speaker: 方彪 (Senior Algorithm Engineer). Team placed 3rd on BIRD Text2SQL benchmark. Talk focuses not on rank but on making capability continuously improve in real business. Answer: turn years of BI semantic layer into AI evolution fuel — low-cost domain evaluation sets, drive effect closed-loop iteration, leverage verified eval sets to boost both accuracy and latency. This year opened CLI for users to complete knowledge precipitation and effect iteration via natural language.

11. 悦点科技 — Above Semantic Layer: Make Enterprise Tacit Experience AI-Executable Assets

Speaker: 任鑫琦 (Founder/CEO). 12 years, ~100 enterprises (manufacturing, auto, energy). Core distinction: domain ontology ≠ renamed warehouse semantic layer; it's an independent enterprise business world model organizing business objects, relations, rules, executable actions. Many firms have data platforms + semantic layers + LLM apps, but projects stall: AI can query/answer but cannot enter core business flows because business concepts, metric calibers, constraints, expert judgment logic remain tacit in docs/heads, never delivered to AI. Lessons: inconsistent terminology/calibers; "big and comprehensive" ontologies obsolete fastest; business/algorithm/data tri-party cognition misalignment. Methodology: scenario-first lightweight ontology modeling + "LLM-assisted extraction + expert review" closed-loop iteration. Case: manufacturing quality traceability across CRM/MES/ERP/PLM — from multi-person days to single-person 5 minutes, yield up 5‰. Quote: "LLM is expression/interaction layer; ontology carries enterprise unique business knowledge; together they form long-term AI assets."

12. 每日互动 — OntoOS: Ontology-Driven Agentic Data Analysis & Strategy Execution

Speaker: 书酉 (Algorithm Expert/AI Architect). Enterprise data analysis evolving from BI and single-turn NL2SQL to Agentic analysis covering query, diagnosis, decision, execution. Core challenge: not just executable SQL, but continuous accurate, explainable, auditable results under complex models, calibers, rules, permissions. OntoOS: enterprise knowledge & business semantic ontology OS — unified management of business objects, metrics, relations, rules, permissions, Evidence. Agent handles intent understanding, semantic alignment, query planning, SQL/graph query orchestration, result validation, analysis explanation. Deterministic engine handles data calc, permission control, rule execution, audit trails. Beyond query: metric decomposition, dimension drill-down, correlation analysis, business knowledge reasoning → verifiable attribution hypotheses → identify key segments → form content/benefit/channel/timing operational strategies. Strategies human-reviewed → audience selection, task orchestration, delivery. Post-activity: response/conversion/cost data → cohort performance, strategy effect, attribution hypothesis review → feedback to next analysis/strategy generation. Background: 个推 AIBI topped "NL2SQL hardest global leaderboard" — accuracy global 4th, efficiency global 6th.

Synthesis

The 85%→50% gap is not in the model but in the unfinished semantic layer. Practical answers to "can we start before semantic layer is done?": 360 — dual-path routing with confidence scores; Gaode — architecture shift to Skill with on-demand retrieval + self-reflection; Ant — infrastructure approach (100+ models, 5,000+ metrics, define once consume everywhere); Zhihu — most complete path: MetricFlow → knowledge base → lightweight KG (1,400+ metrics, 300+ wide tables, 25,000+ SQL templates, 65% penetration). Final takeaway from 悦点科技: "LLM is expression/interaction layer; ontology carries enterprise unique business knowledge; together they form long-term AI assets."

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LLMSemantic Layerknowledge graphmetadata managementevaluation frameworkNL2SQLOntologyData Agent
DataFunSummit
Written by

DataFunSummit

Official account of the DataFun community, dedicated to sharing big data and AI industry summit news and speaker talks, with regular downloadable resource packs.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.