360's Data Agent: Dual-Path Semantic Layer & Confidence-Driven Evaluation
360's enterprise Data Agent bridges the NL2SQL production gap with a dual-path routing system—AST compilation for mature semantic layers and RAG fallback for sparse metadata—guided by confidence scores that gate automated reporting versus human review, while evolving skills into an ontology-driven knowledge layer and validating via a three-level evaluation flywheel.
Large language models achieve 85% SQL accuracy on academic benchmarks like WikiSQL, Spider, and BIRD, but accuracy collapses to around 50% in real enterprise data warehouses. The gap stems from three factors: hundreds of tables and tens of thousands of columns; business definitions that exist only as tribal knowledge; and the inability to detect when the model is wrong.
360's solution adopts a "semi-finished cold start, run and grow" philosophy centered on a dual-path hybrid routing architecture:
AST Main Path: When the semantic layer is complete, the LLM only performs intent extraction; a deterministic compiler then generates SQL, avoiding the "guessing" path.
Fallback RAG Path: When only tables and documentation are available, a hybrid retrieval system (embedding + keyword + foreign-key graph) locates relevant tables, and the LLM generates SQL as a fallback.
Routing Decision: A PreMatch step combined with a confidence metric automatically selects the appropriate path.
Confidence scoring does more than indicate correctness—it determines which scenarios are safe for production. For example, a product operations analyst studying retention changes after a feature redesign could be misled by a single erroneous query. With confidence thresholds, queries scoring above 0.95 are automatically added to reports, while those below 0.65 trigger a "low confidence, human verification recommended" prompt. This gating mechanism is what allows high-stakes, decision-critical workloads to move beyond demos into production.
The skill system evolves through three stages: Thin API → Thick Skill → Ontology Metadata . Initially, thin APIs expose metric, dimension, and enumeration queries, enabling rapid construction of query, attribution, and prediction skills. High-frequency usage patterns—attribution factors, dimension hierarchies, business thresholds—are then reverse-engineered into ontology metadata. In the final state, skills degrade to pure orchestrators while all knowledge resides in the ontology layer.
Academic benchmarks are insufficient because they compare SQL text rather than query results, and they ignore business definitions, multi-turn dialogue, and hallucination detection. 360 built a three-level evaluation architecture:
L1: Programmatic comparison of result sets.
L2: LLM-as-judge for semantic equivalence.
L3: Human fallback for ambiguous cases.
User feedback (likes) is filtered for noise and automatically sampled into a regression test suite. Nine-dimensional metrics are tracked continuously, forming a data flywheel that drives ongoing improvement.
The system currently supports internal big-data cluster operations, S3 and PoleFS storage operations metric analysis and attribution, as well as cross-departmental use cases including data-center operations, document search, and the membership middle platform.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunTalk
Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
