Amap's Text-to-SQL Accuracy Jump: 50% to 95% via Skill Architecture
Amap's intelligent query product raised accuracy from 50% to 95% by replacing RAG with a Skill-based Agent architecture that uses on-demand knowledge retrieval and self-reflection, backed by semantic layer modeling and a three-part evaluation flywheel.
Background: Enterprise Data Service Bottlenecks
Enterprise data services face two persistent bottlenecks: dashboards are never finished, and long-tail needs remain uncovered. Traditional data retrieval only answers "how many," while businesses need "why" and "what to do."
Amap's Intelligent Query Product (Xiao He)
Amap's intelligent query product, Xiao He, did not simply stack a larger model. Instead, they transformed their years-old BI foundation into an Agent foundation that can route and load domain knowledge. The core is semantic layer modeling, not the model itself.
The Failure That Drove the Pivot
The team learned the hard way through a real production failure.
MVP phase: 39 tables, end-to-end accuracy 89.4% — appeared ready for launch.
Full-scale expansion: 331 tables, end-to-end accuracy plummeted to 50%.
After architecture switch: Moving from RAG + Workflow to Skill architecture brought accuracy back to 95%.
Why RAG Collapses at Scale
RAG has three fundamental defects when scaled:
One-time knowledge retrieval — all knowledge fetched at once, causing noise.
Multi-step information loss — context degrades across steps.
No support for cross-table complex calculations.
The solution is a Skill architecture with progressive disclosure: on-demand retrieval plus AI self-reflection for correction, while the Agent framework uses Skills to ensure user depth, controllability, and extensibility for extended query scenarios.
Semantic Layer: Three Knowledge Sources
The semantic layer draws from three sources:
Metadata (auto-extracted)
Business knowledge (human-curated)
Routing rules
Evaluation Flywheel: Three Components
They built an evaluation flywheel consisting of:
AI-Friendly standard score
NL2SQL Exact Match (EM) and Execution Accuracy (EX)
BIAS end-to-end benchmark aligned with OpenAI Trace Grading
Second Track: Agent in Business Decision Flow
Amap data analysis expert Zhong Yujie presented a complementary system that brings Agents into the business decision loop. The local-life business spans 20+ industries with vastly different metric definitions and operational focuses. Her team adopted a three-layer architecture: semantic governance → attribution diagnosis → management judgment.
Accumulated 66,000 traceable knowledge statements.
198 knowledge consumption paths, 186 reaching direct-service state.
Analysis efficiency improved 64×, business approval rate 100%.
Key engineering details include a four-stage knowledge pipeline: edit → compile → publish → consume, and the principle that "numbers are computed by the engine, the model only interprets."
Takeaways
A reusable methodology: transform existing data warehouse/BI assets into an Agent-consumable foundation instead of rebuilding from scratch.
A complete architecture migration story: the decisions behind the 89.4% → 50% → 95% accuracy journey.
Perspective on the ongoing shift from data Q&A (L2) to insight generation (L3), with open challenges.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunTalk
Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
