Ontology Implementation Fork: Semantic Mapping vs Object Storage at Tianyi Credit
The article explores the critical architectural choice between semantic mapping and materialized object storage for ontology implementation, detailing Tianyi Credit's lessons learned, a hybrid approach with core objects materialized, consistency via incremental updates, and a measurable delivery template.
The Fork in Ontology Implementation: Mapping vs Materialized Storage
Ontology projects are widely discussed in enterprises but few succeed in production. Tianyi Credit encountered a recurring dilemma: should the ontology serve as a semantic translation layer over existing data warehouses, or as a materialized object store holding instances, properties, and relationships?
Two Paths, Different Costs
Semantic Mapping Only: The ontology acts as metadata; queries are translated to SQL against the underlying data warehouse. Data stays in place, enabling fast rollout and zero redundancy. However, multi-hop traversals perform poorly, query semantics are constrained by the underlying table structures, and there is no place to persist object state or execution results.
Materialized Object Store: Instances, attributes, and relationships are physically stored. Queries are fast, state can be maintained, and decision loops can be closed. The trade-offs include real-time synchronization, incremental consistency, storage costs, and the "single source of truth" problem.
Irreversible Choice and Lessons Learned
Xu Dehua's team initially chose the lightweight mapping approach. When requirements for multi-hop traversal and state write-back emerged, they were forced to rebuild from scratch. This highlighted that the choice is essentially irreversible once the system scales.
Separating Conceptual and Instance Layers
A root cause of failure is mixing conceptual definitions (object types, properties, relationships, business rules, decision logic) with runtime instance data in the same layer. At scale, this conflation leads to chaos. The adopted separation:
Conceptual Layer: Defines objects, attributes, relationships, terminology, business rules, and decision logic.
Instance Layer: Solves three problems only: queryability, low latency, and accuracy. It handles hundreds of millions of instances, continuously changing states, historical lineages, and result write-backs.
Hybrid Approach: Core Materialized, Long-Tail Mapped
The hybrid strategy materializes core objects while keeping long-tail entities as mappings. The boundary must be fixed upfront, not decided ad hoc. Selection is driven by four criteria:
Query patterns
Latency requirements
Need for state persistence and write-back
Change frequency
Both sides share a unified terminology governance.
Consistency: Incremental Updates with Explicit Uncertainty
A hard constraint: when materialized results diverge from source data, the query side must not return stale data silently. The solution uses incremental updates coupled with invalidation propagation. Before convergence completes, affected query ranges must explicitly label their terminology status — preferring explicit uncertainty over silent stale returns. Otherwise, downstream agents may produce confident but incorrect conclusions.
Data Supply for Agents and Cross-Domain Reuse
Agent-Facing Data: Provide minimal necessary context, governed access interfaces, fallback conditions, and risk metrics.
Cross-Domain Reuse: Establish core vs. extension layer standards and a unified delivery template. The benchmark for success is the reduction in delivery effort for the Nth domain.
Measurable Delivery Template
The team defines a "Business Domain Asset Delivery Five-Piece Set":
Ontology fragment
Rule base
Case library
Verified queries
Evaluation set
Measurement principle: every positive metric must be paired with a cost or safety counter-metric.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunSummit
Official account of the DataFun community, dedicated to sharing big data and AI industry summit news and speaker talks, with regular downloadable resource packs.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
