From 4 Hours to 40 Minutes: Postal Savings Bank's Cross-Unit Data Cleaning Breakthrough

Postal Savings Bank developed a parameter-driven, group-concurrent cross-unit data cleaning framework for its distributed core remittance system, reducing cleaning time for hundreds of millions of records from 4 hours to 40 minutes while ensuring consistency and minimizing business impact.

BanTech Think Tank
BanTech Think Tank
BanTech Think Tank
From 4 Hours to 40 Minutes: Postal Savings Bank's Cross-Unit Data Cleaning Breakthrough

Challenges in Distributed Data Cleaning

As banks complete core system transformation to unit-based distributed architectures, they gain elastic scaling and fault isolation but face data cleaning complexities. Traditional cleaning modes fail to address three key challenges:

Resource pressure: Massive historical transaction and image data cause exponential storage cost growth, degrading query performance and online customer experience.

Cross-unit coordination difficulty: Data scattered across service units creates "data islands," preventing unified scheduling and progress alignment, leading to state fragmentation, broken traceability, and call anomalies.

Efficiency vs. consistency trade-off: Long cleaning durations consume database resources; distributed transaction management complicates checkpoint restart, rollback, and validation, jeopardizing data consistency.

Cross-Unit Cleaning Methodology and Implementation

The bank designed a parameter-driven, group-concurrent flexible cleaning system, turning deletion into a configurable, orchestrated engineering process.

1. Parameter-Driven Intelligent Orchestration

A data cleaning parameter table parameterizes rules, retention periods, operation types, etc. (see Table 1). No code changes needed; only configuration updates define cleaning strategies for any table, decoupling code from business rules.

Table 1: Partial field descriptions of the data cleaning parameter table
Table 1: Partial field descriptions of the data cleaning parameter table

2. Group-Based Cleaning to Eliminate Data Islands

Tables across service units are logically aggregated into business cleaning groups based on business logic. This enables unified scheduling of rules and progress across nodes, ensuring global consistency and alignment, resolving data fragmentation and state inconsistency.

3. Concurrency Control as an Efficiency Multiplier

Dynamic sharding and concurrency control based on configurable parameters. For a service unit with 64 physical shards, setting concurrency to 8 aggregates shards into 8 concurrent batches via load balancing (see Figure 2). This avoids single-node serial bottlenecks, enabling linear throughput scaling. Decoupling shards from execution allows fault tolerance: if a node fails, others take over unfinished shards. Concurrency parameter is tunable based on CPU, memory, and DB IO load.

Figure 2: Dynamic sharding and concurrency control diagram
Figure 2: Dynamic sharding and concurrency control diagram

4. Deep Coupling and End-to-End Tracing for Execution Closure

Concurrency strategy binds deeply with sharding rules: high concurrency for sharded tables, single-thread for non-sharded tables. Supports checkpoint restart and state persistence via full execution logs, enabling seamless re-run and continuation, ensuring full traceability and eliminating data risks from abnormal interruptions.

Practical Results

Deployed in the next-gen core banking remittance system, validated in production. New cleaning tasks respond agilely via parameter configuration, leveraging group aggregation for cross-unit coordination. Under absolute consistency, dynamic concurrency delivers multi-fold efficiency gains over traditional single-threaded mode. For hundreds of millions of records, cleaning time dropped from 4 hours to 40 minutes, avoiding business peak impact. Database disk space precisely released, transaction response stabilized, overall system resilience significantly enhanced. This provides a reusable paradigm for large-scale data governance in banking.

Future Outlook

Using parameter configuration and group concurrency, the bank solved data cleaning challenges in unit-based distributed architecture, proven in production. Data governance is foundational to digital transformation; this cross-unit cleaning technology offers a new technical paradigm for the industry. Future work will continue optimizing data governance to safeguard core system stability and business growth.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

distributed systemsshardingconcurrency controldata cleaningdata governancecore bankingparameter-driven
BanTech Think Tank
Written by

BanTech Think Tank

Tracks major fintech trends, focusing on fintech management, technology development, IT operations, information security, indigenous innovation, data governance, and business innovation. Aims to promote integrated industry‑academia‑research‑application development, offering a sharing platform for tech practitioners and valuable insights for institutional decision‑makers.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.