Building High-Quality Data Sets for Commercial Bank AI: Exploration and Practice
The article examines how commercial banks, driven by national policies and AI advances, develop high‑quality data sets—detailing their strategic importance, regulatory background, challenges in systematization, engineering, and operation, and showcasing ICBC’s layered data architecture, knowledge‑engine pipelines, and future data‑driven competitiveness.
Why High-Quality Data Sets Matter for Bank AI
Amid a new wave of AI breakthroughs exemplified by large‑model platforms such as DeepSeek, data has become one of the three core AI pillars (compute, algorithms, data). The Chinese government now treats the construction of high‑quality data sets as a strategic project for the digital economy, especially as data‑factor market reforms deepen.
Definition and Policy Landscape
In December 2024, the "Guiding Opinions on Promoting High‑Quality Development of the Data Industry" introduced the term “high‑quality data set.” The subsequent national standard, the "High‑Quality Data Set Construction Guide," defines it as a collection of data that, after acquisition and processing, can be directly used to develop and train AI models and effectively improve model performance.
Key policy milestones include:
December 2023: The "Data‑Factor × Three‑Year Action Plan (2024‑2026)" issued by 17 ministries, calling for high‑quality AI training data sets.
December 2024: The State Development and Reform Commission and the National Data Administration formally name “high‑quality data set” as the core carrier for AI‑economy integration.
February 2025: A high‑level kickoff meeting involving 27 ministries clarifies the direction for building such data sets.
Industry Momentum
Financial institutions, especially commercial banks, are accelerating the construction of high‑quality data sets to comply with national directives. In June 2025, the State‑Owned Assets Supervision and Administration Commission convened a summit on AI‑enabled high‑quality data sets for central‑state enterprises, emphasizing the need for standardized, scenario‑driven, intelligent data assets to boost risk control, systemic risk early warning, and personalized services.
Key Challenges for Banks
Incomplete System Architecture: Banking data is fragmented across heterogeneous sources and stored in dispersed silos, lacking a unified management layer.
Low Engineering Automation: Data collection, cleaning, labeling, and quality inspection still rely heavily on manual effort, resulting in low efficiency and high cost.
Complex Operations: High‑quality data sets must cover structured, semi‑structured, and unstructured modalities while meeting strict accuracy, completeness, consistency, and compliance standards, requiring coordination among business experts, data engineers, AI engineers, and security engineers.
ICBC’s Exploration and Practice
Industrial and Commercial Bank of China (ICBC) has pioneered an enterprise‑level data‑element engineering project. By leveraging an enterprise data‑mid‑platform, ICBC built a three‑layer data hierarchy—"source layer → aggregation layer → extraction layer"—to enable large‑scale structured data storage, sharing, and AI application support.
For unstructured data, ICBC created an end‑to‑end knowledge‑engine pipeline that integrates data acquisition, cleaning, annotation, quality assessment, and operation, forming a robust knowledge‑operation capability for AI model training.
1+4+X High‑Quality Data Set Framework
ICBC adopts a "1+4+X" model:
1: An enterprise knowledge base that aggregates business knowledge data for unified management.
4: Four categories of training data—pre‑training, fine‑tuning, chain‑of‑thought, and evaluation datasets.
X: Inference data focused on real‑world model application and staff empowerment.
This structure ensures that AI models understand both financial domain expertise and complex, evolving customer needs.
2) One‑Stop Knowledge Engineering Pipeline
ICBC built an online, centralized, full‑link engineering pipeline covering data collection, processing, management, and application. Intelligent technologies such as AI‑assisted labeling and automated quality inspection are embedded to guarantee efficiency and accuracy. The resulting "data flywheel" continuously enhances intelligent finance capabilities.
Data‑asset achievements include a multi‑TB financial training data set and over 8 million knowledge items. The high‑quality data set now spans more than 20 core business domains, and the knowledge‑Q&A service has handled over 4 billion queries with an accuracy exceeding 90 %.
3) Operational Mechanism
ICBC follows a six‑step process for high‑quality data set construction: planning & design → admission assessment → labeling & quality inspection → model verification → system deployment → scenario operation. Standardized data‑quality, cleaning, desensitization, and annotation guidelines enforce consistency.
Collaboration roles are clearly defined (see Figure 1): business experts handle planning and data review; application developers verify technical feasibility; data engineers manage data processing and model optimization; product managers analyze business needs; data scientists conduct admission assessment and model training.
Future Outlook
With deeper data‑factor reforms and widespread AI adoption, commercial banks will merge internal and external data sources to build a unified, lifecycle‑covering data system. Technologies such as privacy‑preserving computation and automated annotation will lower processing costs and enhance security. New regulations that require data assets to be recorded on‑balance will turn high‑quality data sets into core assets, spurring data‑trading models and cross‑industry data‑space collaborations. Ultimately, a bank’s competitive edge will hinge on its data‑driven capabilities, using continuously optimized high‑quality data sets to achieve real‑time risk control, personalized services, and intelligent operations.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
BanTech Think Tank
Tracks major fintech trends, focusing on fintech management, technology development, IT operations, information security, indigenous innovation, data governance, and business innovation. Aims to promote integrated industry‑academia‑research‑application development, offering a sharing platform for tech practitioners and valuable insights for institutional decision‑makers.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
