How Ant Group Scaled Apache Ossie Semantic Layer from Zero to 5,000 Metrics
The article explains why large language models struggle with business data, introduces Ant Group's semantic‑layer approach built on Apache Ossie to impose strong business constraints, compares it with other retrieval methods, and details the engineering journey that delivered a unified source of truth for over 5,000 metrics while also announcing a related conference.
Large language models often produce hallucinations, unstable answers, and lack confidence when directly answering business data questions; for example, asking for last month’s East China GMV may yield fabricated numbers, inconsistent responses, or ignorance of changing definitions.
Ant Group addresses this gap by inserting a semantic layer between the model and raw data. The layer structures all business knowledge—metrics, dimensions, entities, processes, physical implementations, and SQL rules—so the AI operates within explicit constraints, generates SQL, and explains results instead of touching naked data.
Compared with text knowledge bases, RAG vector retrieval, and full‑metadata retrieval, the semantic layer offers three sharper advantages: clearer query constraints, explainable outcomes, and lower inference cost. It serves both human analysts and AI agents, delivering accurate answers for both.
After platformization, Ant unified more than 100 models and over 5,000 metrics under a single source of truth. The semantic layer also resolved long‑standing issues such as expired definitions, downstream desynchronization after table changes, inconsistent local versions, and scattered knowledge across teams by centrally managing definitions, versions, mapping relationships, and consumption governance.
The foundation is Apache Ossie, an open‑source semantic‑layer specification of which Ant is a core contributor. The presentation walks through the evolution from zero to 5,000 metrics, covering semantic modeling, semantic retrieval, ontology construction, semantic mapping, evaluation framework, MCP integration, and data‑governance practices—all in a single engineering pipeline.
For further details, the talk was delivered by Wang Xiaojun at DACon Beijing on 23‑24 Oct 2026, with a limited‑time 20% discount and registration via the provided link.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunSummit
Official account of the DataFun community, dedicated to sharing big data and AI industry summit news and speaker talks, with regular downloadable resource packs.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
