10 Financial Firms Share AI Agent Strategies for 98.5% Auto-Review, 0.003% Fraud
This article analyzes how 10 leading financial institutions implement AI agents in low-tolerance scenarios, detailing their approaches to data ontology, semantic layers, multi-agent architectures, risk control, and evaluation frameworks, achieving metrics like 98.5% automated review rates and 0.003% fraud rates while ensuring auditability and regulatory compliance.
Three Fundamental Barriers to AI in Finance
The article opens by establishing why finance sits at the top of the AI adoption difficulty curve: AI outputs here are not advisory opinions but binding decisions backed by regulation, audit, and real money. Three barriers emerge from practitioner quotes:
Data Wall : Fragmented, offline, non-standardized data (contracts, logistics, customs, invoices at XTransfer; research reports, announcements, roadshows at Tianhong Fund; inconsistent definitions of "active user" across business lines at Ant Group). As XTransfer's Chen Wensong notes, "data unusable" often precedes "model not strong enough" as the first blocker.
Model Hurdle : Finance demands not "roughly right" but "right every time" with full explainability. Mashang Finance's Zhao Xuebo cites long Agent build times, unstable generation, high resource costs, and — critically — lack of optimization and evaluation drivers for continuous improvement. Airwallex's Dong Daifan adds that individual Agent accuracy does not guarantee safety of the combined multi-Agent workflow.
Responsibility Gate : Unique to finance. Airwallex classifies capabilities into Read, Analyze, Propose, Act levels — only the first two can run automatically within authorization; high-risk actions (modifying cases, sending RFIs, changing workflow state) require deterministic permission checks and human approval. Dun & Bradstreet retains manual review and audit trails; XTransfer's principle is "AI assists decisions, does not replace decisions" with traceability, auditability, and human intervention.
Layer 1: Making AI Understand Financial Business — Getting the "Definitions" Right
01 | Ant Group: Financial Data Knowledge Ontology — From "Querying Data" to "Understanding Business"
Speaker : Cui Weibin, Senior Data Technology Expert, Ant Group.
Core Problem : Business lines define "active user" differently; source tables are massive; manual governance is unsustainable.
Approach : Build a Data Knowledge Ontology to unify multi-source business semantics. Instead of forcing a single definition, the ontology preserves differences and lets Agents fetch on demand. Technical pipeline uses LangGraph six-node state machine (
InputParser → Comprehender → EntityResolver → Reconciler → Synthesizer → Verifier) for full-table perception, entity recognition, cross-table conflict resolution, ontology synthesis and validation. SQL Understander (sqlglot + regex) reverse-extracts field semantics and table relationships from historical SQL code — "code as annotation" for automated ontology generation. Service layer uses a mutually-indexed architecture with layered joint retrieval to serve Agent analysis requests.
Three Knowledge Layers Retrieved Synchronously : KG_cs (business perspective, e.g., Huabei "active user" definition rules) KG_fr (entity-relationship graph, e.g., "User A — holds — Huabei limit 50k") RC (raw document evidence)
Results are fused, deduplicated, ranked, and packaged as a Knowledge Package .
Key Stance : "Ontology is not a dictionary; it is the textbook that lets the system 'understand business'."
Results : Covers thousands of tables; cross-selling opportunity identification reduced from 3–5 days to hours; anomaly root-cause analysis accuracy >85%; supports multiple marketing strategies in production.
Outstanding Challenges : 7-level conflict resolution achieves 60–70% auto-merge accuracy in extreme semantic divergence; 20–30% still need human intervention, with linear cost growth as tables scale. Current system is passive (user-triggered); next step is proactive business-change detection and insight push, requiring stronger causal logic in KG_cs.
02 | Ant Group: Apache Ossie Semantic Layer — From 0 to 5,000+ Metrics
Speaker : Wang Xiaojun, Head of Ant Data Semantic Layer Platform.
Core Insight : Data Agent bottleneck has shifted from "SQL generation" to "understanding business semantics"; semantic layer is the foundation for accuracy, trust, and scale.
Solution : Based on open-source OSI, adopt "Base Standard + Business Extensions" . Asset platform carries a Single Source of Truth (SSOT) for unified models, lineage, and consumption outlets — "define once, consume everywhere".
Progress : 100+ semantic models, 5,000+ metrics; multi-scenario consumption for humans, Agents, and platforms; deployed across multiple internal business scenarios.
Challenges & Solutions :
Native OSI lacks support for multi-definition metrics, multi-granularity aggregation, equivalent metric definitions, public dimension reuse, cross-model references → built controlled extension standard supporting complex metrics, semantic views, model composition, and AI context expression.
Models drift with dev/business changes; manual maintenance costly and outside dev loop → full-chain monitoring (feedback, patrol, change analysis) + forward dev loop (embedding semantic model iteration into requirements, development, release, consumption) to reduce freshness cost.
Agent struggles to locate correct context among thousands of similar/related semantic objects; full loading causes context bloat and inference latency → "Hybrid Retrieval + Progressive Disclosure" consumption mechanism combining keyword, vector, business tag, and graph relationship recall with confidence scoring and re-ranking.
03 | Tianhong Fund: Turning Research Reports, Announcements, Roadshow Audio from "Files" into "Data Objects"
Speaker : Tian Jilong, Senior Data Engineer, AI & Data Dept, Tianhong Fund.
Context : Fund companies accumulate massive unstructured investment-research data daily (reports, announcements, roadshows, public sentiment, expert interviews) — sources scattered, formats complex, standards inconsistent, long untreated and uncomputable.
Engineering Pipeline :
Collection & Access → Document Parsing → Content Cleaning → Unified Modeling → Data Governance. Further uses NLP & LLM as new semantic processing capabilities for entity, event, viewpoint extraction, semantic chunking, vectorization — gradually converting raw documents into governable, linkable, computable data assets .
Key Stance : "Don't treat LLMs as simple Q&A tools; treat them as a new semantic computation and data processing engine, combined with traditional ETL, rules, NLP, and data warehouses to automate previously intractable data tasks."
Challenges : Document structure fidelity vs. chunk granularity trade-off; retrieval precision and result trustworthiness in financial scenarios.
Layer 2: Low-Tolerance Scenarios — From "Usable" to "Trustworthy"
04 | XTransfer: B2B Cross-Border Financial Risk Control — Three-Step AI Productization
Speaker : Chen Wensong, Senior Product Director, XTransfer.
Scene : B2B cross-border trade risk control — high-value low-frequency transactions, multi-party document cross-verification, multi-jurisdiction compliance. General LLMs and traditional rule engines both fall short.
Three-Step Methodology :
Data → Knowledge : Build B2B Trade Knowledge Graph (cross-border industry knowledge base + enterprise transaction graph database); key is standardization and semanticization of multi-source heterogeneous data.
Recognition → Decision : Self-trained vertical multi-modal model TradePilot for multi-source document structured extraction, visual forgery/tampering detection, cross-verification of transaction rationality.
Usable → Trustworthy : Agent testing, sandbox simulation, A/B testing, gray-scale dual-run, online evaluation & monitoring.
Three Laws :
Law 1: Trustworthy Data > Powerful Model . Moat is high-quality, traceable domain data.
Law 2: Trustworthy Decision > Flashy Capability . Partners, regulators must understand, trust, audit.
Law 3: Trustworthy Value > Concept Leadership . AI product must answer: how much cost saved, risk reduced, efficiency gained.
Production Results : Automated review rate 98.5% , fraud rate 0.003% , risk control cost reduced 20%+ .
Engineering Balance : "Rule Engine + AI Model + Human Expert" collaboration — AI discovers risk signals, rule engine guards explainable hard bottom line, humans handle only AI-flagged gray zones.
05 | Airwallex: Financial Risk Control Multi-Agent — How to Achieve "No Overreach"
Speaker : Dong Daifan, Risk Architect, Airwallex.
Architecture Evolution : From single-case Case Copilot to cross-case Agent Mesh . Each case system gets an independent Domain Agent retaining its own facts, rules, and decision ownership. Agents expose governed capabilities via A2A Contract . A Cross-Case Orchestration Layer handles capability discovery, task tracking, evidence linking, conflict detection, and cross-case analysis.
Core Design Principle : "Production orchestration is not handed to a fully autonomous 'brain'; instead deterministic workflow + bounded LLM planner : deterministic code handles permissions, state, retries, compensation, and high-risk actions; LLM only handles fuzzy intent decomposition and constrained dynamic routing."
Eight Production Questions Answered :
Agent understands business : Domain knowledge generated from real sources (code, API, rule configs, SOP, expert experience), reviewed and versioned by domain owners, distilled into testable, reusable skills .
Huge prompts cause instability : Decompose full judgment into independent checkpoint skills — each capability separately developed, evaluated, cached, deployed, rolled back.
Long tasks & timeouts : Gateway timeout ≠ downstream task not executed; naive retry causes duplicate ops. Use A2A Task lifecycle management with unique Task ID, idempotency key, async status query, artifact persistence for traceability, recoverability, deduplication.
Cross-case fact conflicts : Different case systems have inconsistent data freshness, rule versions, business semantics. Orchestration layer forces each Agent to return data source, observation time, version info ; inconsistencies generate explicit conflict artifacts routed to correct domain owner.
Permission & responsibility boundaries : Four levels — Read, Analyze, Propose, Act. Read/Analyze auto-execute within authorization; high-risk actions (modify case, send RFI, change workflow state) require deterministic permission checks and human approval.
Multi-Agent error accumulation : Single Agent accuracy ≠ combined workflow safety. Beyond per-Agent benchmarks, evaluate end-to-end workflow: human baseline, material error, evidence completeness, routing accuracy, unsafe action rate.
Safe failure : On Agent timeout, unavailability, schema incompatibility, orchestration layer returns explicit partial result stating what completed and what is temporarily unavailable — never let the model hallucinate missing info . Also requires circuit breaker, capability kill switch, version rollback, human-process fallback.
Serving different roles : Risk needs full evidence, Ops needs next action, CS cares who handles and what can be told to customer, management needs cross-case trends. Solution: generate role-specific outputs on shared evidence base.
06 | Dun & Bradstreet: When Agent Enters Regulatory Process — Cross-Border KYB & UBO Identification
Speaker : Feng Zhikai, China Product Director, Dun & Bradstreet.
Regulatory Driver : Deepening AML Law and PBOC requirements for Ultimate Beneficial Owner (UBO) identification in cross-border account opening, merchant onboarding, trade finance, cross-border payments. Traditional KYB relies heavily on manual search and verification — efficiency and consistency challenged.
Architecture : Data → MCP → Skills → Agent → Business Closed Loop. Core Stance : "AI improves efficiency; humans make decisions."
Key Practices : Introduce trusted enterprise identity system (D-U-N-S® Number) , global corporate relationships & UBO data at Agent base layer. Standardize enterprise verification, equity penetration, risk assessment via MCP and business Skills. Retain manual review and audit trail mechanisms .
07 | Mashang Finance: Turning "Delivery Rate" into an Engineering Object
Speaker : Zhao Xuebo, Big Data Director, Mashang Finance.
Pain Points : Analysis/decision Agent build time long, generation unstable; high resource cost & latency; optimization & evaluation insufficient — no driver for continuous improvement .
Three-Step Solution :
Classify Agent analysis decisions, differentiate element lists per decision type.
Build knowledge on decision elements (industry, company, domain, scenario, technology) to equip Agent with effective tool-chain planning and precise prompt context.
Establish two-level scoring system for Agent decisions, digitize management to form improvement drivers.
Challenges : How to decompose Agent execution chain to atomic steps for single-variable observability/improvement; how to measure objective digital Agent effects consistent with business user experience.
Layer 3: Platformizing Capabilities — Engineering, Cost, Evolution
08 | Lingyue: Skill vs. Agent Boundary — Where to Fix a Wrong SQL
Speaker : Guo Zhihao, Data Application Lead, Lingyue.
Entry Point : A SQL parses, executes, returns results — yet yields wrong business conclusion.
Core Model: "Four Layers + One Plane" Responsibility Division :
Knowledge & Context Layer : Versioned management of business definitions & table/field semantics; on-demand retrieval, assembly, update of Context during task.
Skill Method Layer : Versioned, reusable methods for SQL generation, probing, validation.
Workflow & Agent Decision Layer : Workflow handles codable states/branches; Agent only handles local judgments that must combine user input and tool feedback.
Tool Execution Layer : Structured capability interfaces — metadata query, restricted data probing, SQL validation.
Horizontal Governance Plane : External systems enforce identity/permissions, parameter policies, budgets, approvals, termination; observability & evaluation provide cross-layer evidence and regression verification.
Autonomy Decision Table : Agent candidate must satisfy — clear goal & success criteria, semantic judgments uneconomical to code, timely trustworthy feedback available, action space constrainable, result verifiable. Then use risk, verifiability, reversibility to set autonomy level.
Live Review Examples : Data quality checks → deterministic Workflow; SQL generation → constrained Agent only in clarification, probing, generation steps; final execution still platform-controlled & user-confirmed.
Key Stance : "Let Agent guess less, let platform know more."
09 | China Telecom Payment: Ontology Implementation Fork — An Irreversible Choice
Speaker : Xu Dehua, GM Assistant, Risk Management Dept Director, Model Team Lead, China Telecom Payment / China Telecom Credit.
Core Dilemma : Is ontology a "translation layer" (semantic mapping only) or a "data layer" (materialized object store)?
Mapping Only : Ontology as metadata; query-time translation to DW SQL; data stays put, fast launch, no redundancy. But multi-hop penetration fails , definitions bound by underlying table structure, object state & execution results have no home .
Build Object Store : Materialize instances, attributes, relationships; fast queries, carries state, supports decision closed-loop. But faces real-time sync, incremental consistency, storage cost, "who is the truth" .
Upstream Insight : Concept definitions and real business data must not be mixed in one layer .
Four Selection Criteria : Query patterns, latency requirements, need for state/write-back, change frequency.
Hybrid Solution : Core objects materialized, long-tail keeps mapping; boundary written in stone, not ad-hoc . Consistency constraint: before change convergence completes, affected query results must explicitly label definition status — "rather explicitly state uncertainty than silently return".
Delivery Template for Business Domains : Ontology fragment, rule library, case library, verified queries, evaluation set. Metric principle: every positive indicator paired with a cost/safety counter-indicator.
10 | Xiaoying Technology: From Buy to Build — Four-Stage Evolution of Conversational Bots
Speaker : Wu Wenbin, Model Application Development Lead, Internet Platform Dept, Xiaoying Technology.
Motivation : Third-party bots suffered from hard customization, high long-term calling cost, shallow business process integration .
Four-Stage Evolution :
Vector Knowledge Base + RAG basic dialogue.
LLM integration, standardized SOP process bots.
Upgrade to Agentic architecture, autonomous task-planning voice bots.
(Ongoing) Multi-agent collaboration, continuous inference cost reduction, automated evaluation system.
Key Technical Breakthroughs : Solved real-time voice link latency, Agent process runaway, knowledge base hallucination ; built layered memory & task scheduling system . Post-launch: significant uplift in human-machine dialogue completion rate, long-term service cost reduction.
Challenges & Solutions :
Agent execution uncertainty : Complex flows cause SOP deviation, hallucination. Fix: soft+hard rule dual constraints, task state persistence, multi-round behavior verification.
Voice link + LLM joint inference latency : Audio streaming, vector retrieval, serial LLM calls cause timeouts. Fix: module async decoupling, hot knowledge pre-caching, tiered inference strategy.
Cross-Cutting Pattern: No One Talks About Model Strength
Across all 10 sessions, not a single presenter leads with model parameters or benchmark scores. The shared focus is on engineering the guardrails :
How to handle definition conflicts (Ant Group)
How to keep semantic models fresh (Ant Group)
How to turn research reports into computable assets (Tianhong Fund)
How to turn trade documents into trusted evidence (XTransfer)
How to tier permissions (Airwallex)
How to penetrate UBO structures (Dun & Bradstreet)
How to build scoring systems (Mashang Finance)
How to draw boundaries (Lingyue)
How to choose implementation paths (China Telecom Payment)
How to evolve from buy to build (Xiaoying Technology)
As Chen Wensong puts it: "The moat of financial AI risk control is not the algorithm; it is high-quality, traceable domain data."
And Dong Daifan's pattern may be the most copy-worthy: "Deterministic code owns permissions, state, retries, compensation, and high-risk actions; LLM only owns fuzzy intent decomposition and constrained dynamic routing."
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunSummit
Official account of the DataFun community, dedicated to sharing big data and AI industry summit news and speaker talks, with regular downloadable resource packs.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
