From 50% to 95%: Amap's 4-Talk Blueprint for Trustworthy Decision Agents
Amap shares four conference talks detailing how they built trustworthy AI agents for business decisions, covering semantic governance, causal evaluation, evaluation-driven development, and a product architecture shift that lifted accuracy from 50% to 95% by separating deterministic computation from model reasoning.
Overview: Four Talks Form a Complete Engineering Chain for AI Agents in Business Decision-Making
Amap (Gaode) presented four sessions that together address three systemic breakpoints: data disconnected from business context, diagnostic conclusions disconnected from evidence, and analysis reports disconnected from management actions. The four talks — semantic governance, causal evaluation, evaluation-driven development, and product architecture migration — each tackle one link, converging on the principle: code owns facts, numbers, identity, versioning, and delivery state; models own limited business judgment and natural-language expression.
Talk 1: Semantic Governance, Attribution Diagnosis & Management Judgment (Zhong Yujie)
Problem
Local-life business spans 20+ industries with vastly different metrics, definitions, and operating focuses. Manual analysis is slow and experience is not reusable. Three breakpoints appear: data–context disconnect, conclusion–evidence disconnect, report–action disconnect.
Architecture: Three-Layer Agent
Semantic Governance Layer — unifies business definitions and knowledge via an "edit–compile–publish–consume" pipeline. Published versions are immutable (content-hash fixed); ambiguities, gaps, expirations, and truncations are returned explicitly. Result: ~66,000 traceable knowledge statements, 198 consumption paths (186 directly servable).
Attribution Diagnosis Layer — closes metric identities first, then runs anomaly detection, LMDI contribution decomposition, trend analysis, and dimension drill-down. Numbers and attribution results are computed by a deterministic engine; the model only interprets, never generates or rewrites numbers.
Management Judgment Layer — scans all operating units across a full cycle, distinguishes management issues from normal states; normal states enter strategy review to avoid manufacturing problems . Conclusions are bound to evidence, guarded, and equipped with failure-circuit-breakers and retrospective conditions for auditability.
Results
New-industry onboarding cut from half-month to 1 day.
Single-industry report delivery closure rate 100% (latest cycle).
Analysis efficiency up 64×; business acceptance 100% (speaker-reported).
Key Challenges & Solutions
Fragmented, evolving semantics: unified governance with scope, version, provenance, and explicit invalidation.
Insights lacking evidence constraints: deterministic engine computes; model does constrained judgment; evidence binding + guards + circuit-breakers + retrospective conditions create auditable decision support.
Talk 2: From One-Off Evaluation to a Trustworthy Causal Engine (Xu Xiaowei)
Problem: Marketing ROI as a Counterfactual
"Did the subsidy bring incremental orders?" is a causal question. The fundamental difficulty: counterfactuals are unobservable — there is no ground truth to verify against, so a model's "always gives an answer" instinct becomes a risk.
Architecture Decision: "Error Observable vs. Error Silent"
They chose PSM + DID over simple pre/post or pure A/B. The critical architectural choice: AI only touches orchestration, never numerical computation. Deterministic engine calculates; model orchestrates.
Trust Engineering in a No-Ground-Truth World
Dual-brain architecture (frozen core + Agent Loop) with parity contracts for reproducibility.
Confidence grading + definition gates + multi-method triangulation — the system learns to say "I'm not qualified to conclude."
Three real incident retrospectives (inflated ROI, false negative increment, false significance) show the gap between "giving an insight" and "being honest about confidence."
Open Challenges
~15% of activities land in a gray zone: "evaluable but evidence insufficient" (L4 fallback).
No agreed offline ranking for conflicting causal methods; exploring Energy Score for automatic arbitration.
Takeaway
In decision systems without ground truth, engineering focus shifts from "pursue higher accuracy" to "manage the system's self-awareness of correctness."
Talk 3: From TDD to EDD — Evaluation-Driven Data Agent Development (Rong Keke)
Why Academic Benchmarks (BIRD/Spider) Fail
Static test sets vs. continuously evolving business definitions.
Single EX metric vs. business-definition correctness ( result correct ≠ definition correct ).
Standard single-db vs. real warehouse scale.
Four-Stage Pipeline: Build – Manage – CI – Attribute
Scoring framework (4 dimensions): Task Completion 45% / Reasoning & Planning 25% / Engineering Efficiency 15% / Safety & Robustness 15%. Top-level metrics: EX, latency. Vertical metrics: scenario routing, metric recall, SQL consistency.
Build: 1,000 questions from real online issues + lab questions, allocated across 4 business directions, stored in a question-bank platform (production-ready).
Manage: 14-day rolling updates + old/new version snapshots; automatic detection of definition and table changes.
CI: Parallel scheduling, checkpoint resume, answer-lock isolation → unattended evaluation. Cycle reduced 24h → 12h; human effort down ~30%.
Attribute: Aggregation conservation checks, L1/L2 routing hit-rate, SQL consistency. Misclassification → error taxonomy → feeds back into Skill improvement.
Safety evolution: Prompt guidance → sandbox interception → hard isolation (read-only credentials + ACL).
Open Challenges
Evaluator's own trustworthiness: golden answers drift with definitions; machine judgments have false negatives (e.g., aggregation granularity mismatch). Still needs human spot-checks; clean separation of "Skill true error rate" vs. "evaluation undecidable rate" unsolved.
Loop from "find error" to "fix error" not closed: attribution still manual; definition knowledge scattered in personal experience, not auto-structured to feed Skill generation.
Talk 4: From "Answers Right" to "Insights Visible" — BI-to-Agent Base Migration (Lin Yuhang)
Context & Gartner Ladder
Enterprise data service stuck: dashboards never finish, long-tail needs uncovered; traditional query only answers "how much", not "why/what next". Positioned at Gartner L2 (intelligent Q&A) moving to L3 (intelligent diagnosis); L4 (autonomous decision) still blank, 12-18 month window.
Two-Engine Strategy
Xiao-Hu ChatBI — universal L2.
Cici Skill + Agent — deep L3/L4.
Core is semantic-layer modeling, not the model itself .
Real Two-Month Migration: Three Versions
MVP (39 tables): 89.4% accuracy.
RAG + Workflow full scale (331 tables): end-to-end accuracy collapsed to 50% . Root causes: one-shot knowledge retrieval, multi-step information loss, no support for cross-table complex calculations.
Skill Architecture (progressive disclosure): on-demand retrieval + AI self-reflection correction → accuracy recovered to 95% . Skill architecture also gives controllability and extensibility for user depth and extended use-cases.
AI-Friendly Knowledge Base (Three Sources)
Metadata (auto-extracted)
Business knowledge (human-fed)
Routing rules
Evaluation Flywheel
AI-Friendly standard score
NL2SQL EM + EX
BIAS end-to-end benchmark against OpenAI Trace Grading
Industry Benchmarks
ByteDance / Meituan NL2SQL 95%; OpenAI 6-layer context — used for calibration.
Ongoing Challenges
L3 insight credibility, user trust, information loss.
Scenario coverage 80% → 95%, response speed, L3 scaling.
Synthesis: The Dependency Chain
Definitions not governed → attribution numbers meaningless. No confidence judgment → system forces conclusions in no-ground-truth scenarios. No evaluation → you don't know if it's accurate, nor whether a change improved or regressed. First three undone → product remains a demo.
The four talks are not independent case studies; they are four gates on a single path. Amap also openly lists unsolved problems: evaluator trustworthiness, the ~15% gray-zone activities, L3 user trust, and automatic feedback of definition knowledge into Skills.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunSummit
Official account of the DataFun community, dedicated to sharing big data and AI industry summit news and speaker talks, with regular downloadable resource packs.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
