Tencent's 32B Financial Compliance Agent: Skill Rules & Dual-Track Self-Evolution

Tencent's financial compliance agent uses a 32B model, skill-based rules, and dual-track self-evolution to achieve 95%+ recall, 80%+ precision, 40-second inference, and hourly rule updates, replacing manual review that took days.

DataFunTalk
DataFunTalk
DataFunTalk
Tencent's 32B Financial Compliance Agent: Skill Rules & Dual-Track Self-Evolution

Why Manual Review Fails

Financial marketing compliance review is not just checking copy. Two regulatory cases illustrate the problem: process failures (posters launched without proper compliance checks, missing risk disclosures and personnel qualifications) and selective disclosure (showing only partial strong performance). Both point to the same issue: a single compliance decision must verify content, data, entity, and process simultaneously, forming a complete evidence chain .

Review objects have expanded from text-only (customer service messages, app pop-ups) to posters, banners, long product pages, short videos, and live streams. The traditional linear workflow—receive materials, queue dispatch, human judgment, spot-check/review—means material volume growth directly demands proportional headcount growth . This leads to insufficient coverage, review cycles measured in days or weeks, and inconsistent judgments across reviewers. New regulations also create a window where old practices persist.

The system was redefined with four goals: see everything, judge fast, judge accurately, keep up —unifying multi-material intake, pre-publish review, stable re-review, and rapid new-rule effectiveness into system capabilities.

Turning Regulations into Agent-Executable Rules: From Long-Image Parsing to Evidence Chain

Online Pipeline Architecture

The online chain: incoming materials → multimodal parsing + entity/scene identification → skill rule routing, risk reasoning, evidence anchoring, reverse verification → output risk tags, levels, violation evidence, cited clauses, and modification suggestions. Supporting capabilities include perception parsing, tool calling, rules & strategies, compliance reasoning, knowledge & memory. Offline handles skill governance, rule self-evolution, and model data flywheel.

Perception Parsing: Long-Image Parsing Challenges & Solution

Fund promotion long pages contain product info, performance charts, manager bios, and bottom risk disclosures—requiring content completeness while preserving visual semantics and spatial relationships. Pure OCR loses charts and positional relations; general multimodal models miss bottom fine print, hallucinate text, or fail to parse complex charts fully.

Engineering solution: three-stage processing

Semantic slicing by visual boundaries, background changes, and OCR text boxes; adjacent slices retain overlap.

Feed each slice to multimodal model to extract text, charts, elements, and positional relations.

Sort by original image coordinates, deduplicate overlaps, and semantically stitch.

Results: recall improved from 90.15% to 97.59% , F1 from 94.73% to 98.67% , miss rate dropped from 9.85% to 2.41% , average latency increased from 16s to 21s.

Long image parsing three-stage pipeline and metrics
Long image parsing three-stage pipeline and metrics

Rule Sources & Structuring

Rule base comprises three layers: basic laws, financial industry special rules, content & AI governance , continuously ingesting regulatory bulletins, penalty cases, and industry self-discipline rules. Each rule carries regulation/file number, original text, applicable entity & scenario, effective status, version, and case basis—ensuring traceability and updatability.

Why Raw Regulations Can't Be Used Directly

Feeding raw regulations to an agent hits five execution breakpoints : content too long/scattered, abstract execution boundaries, unclear applicability scope, missing evidence & disposition structure, and version drift. Corresponding engineering transformation also has five steps:

Map by entity, product, channel, time.

Atomicize clauses into entity, action, object, condition, exception, measurable thresholds.

Build applicability matrix: entity × product × channel × audience × action.

Add decision conditions and human routing.

Validate with positive, negative, and boundary examples before versioned release.

In fund promotion, tag system handles routing, structured rules handle adjudication , covering promotional claims, historical performance, index data, risk disclosure, and licensed business risks.

Skill-Based Governance

As rules grow, stuffing all into the prompt causes linear token cost growth and attention dilution. Adopted two-level progressive disclosure : core risk rule definitions stay in prompt for high-recall initial screening; on risk hit, load the corresponding skill's boundaries, exceptions, and historical cases. Each inference sees only the rule subset relevant to the current material.

Two-level progressive disclosure for skill governance
Two-level progressive disclosure for skill governance

Tool Calling for Auditable Evidence

Compliance conclusions rely on dynamic facts outside the material. Example: "fixed income plus" promotion needs actual holdings data not in the image. System performs entity disambiguation , plans data sources, query fields, and valid time points, calls external APIs via MCP parameterized invocation; results validated for type, unit, range, source → form evidence snapshots with timestamp, tool version, raw response. Timeouts, missing data, or source conflicts escalate to human review.

Tool calling and evidence snapshot generation
Tool calling and evidence snapshot generation

Evidence Anchoring: The Deliverable Is Evidence, Not Just Judgment

System maps conclusions back to original image coordinates: text via character boxes, visual objects via open-vocabulary object boxes, irregular regions via pixel contours—using EasyOCR/CRAFT, Grounding DINO, SAM. Localization failures or evidence conflicts also fall back to human review.

Evidence anchoring with multiple localization techniques
Evidence anchoring with multiple localization techniques

Why Train a Small Model: Financial Compliance Needs Not "General Capability"

Problems with General LLMs

Vertical risk identification accuracy/recall only ~30-40% .

Large models: single inference 128 seconds on 16 GPUs .

General training objectives don't align with compliance's "high recall first" requirement.

Average ~15k token long prompts easily lose focus.

Team trained a task-specific model targeting

high precision/recall, low latency, long-context localization

.

General LLM vs custom model comparison
General LLM vs custom model comparison

Three-Stage Training

Data synthesis (dual-path):

Reverse synthesis: from human-confirmed gold-risk samples → correct to risk-free text → generate new samples per risk label.

Forward synthesis: unlabeled materials → OCR, long-image slicing, multi-model labeling merge → human verification.

SFT : teaches review requirements and structured output, kept lightweight to avoid over-training.

Reward models & RL : based on full compliance quadruples and negative examples (missed, false, hallucinated) construct pair preference data ; train model-based and rule-based reward models.

Training pipeline: data synthesis, SFT, RL
Training pipeline: data synthesis, SFT, RL

Three Key Decisions in SFT/RL Combination

Data coverage: native/synthetic, risky/safe, long-tail categories—avoid head-risk dominance.

Keep SFT light to leave exploration space for RL.

Base model selection: not just Pass@1, but Pass@8 to observe potential ceiling .

Key training decisions
Key training decisions

Business-Adjusted Objective: F2 Score

Standard F1 weights precision and recall equally. Compliance fears misses more, so adopted F2, making recall weight 4× precision .

F2 objective for compliance
F2 objective for compliance

Dual-Track Self-Evolution: Rule Changed Today, Capability Grows Tomorrow

Rule Self-Evolution (Fast Track)

"Self-evolution" does not mean the agent rewrites rules and deploys directly. Inputs: misjudgments/misses from legal review, proactively scanned regulatory signals, rule conflict detection. AI inductively identifies uncovered rules, missing exceptions, fuzzy boundaries and generates suggestions; legal confirms, corrects, or rejects. Confirmed rules enter rule library, next inference dynamically loads them—no model retraining needed. Updates achievable in hours .

Rule self-evolution loop
Rule self-evolution loop

Dual-Track Evolution

Rule track : solves "today's problem fixed today"—legal feedback directly becomes skill updates for rapid correction.

Model track : distills experience into capability—confirmed/corrected/rejected labels become RL signals for periodic GRPO retraining.

Both tracks share the same real feedback source but handle short-cycle rule changes and deep model capability separately.

Dual-track evolution architecture
Dual-track evolution architecture

Results & Conclusions

Coverage: from sampling to full coverage .

Cycle: 3-5 days → minutes .

Custom 32B model: recall 95%+, precision 80%+ .

Inference: 128s (671B) → 40s (32B) , ~69% reduction; deployment compute cost -75% .

Rule updates: weekly+ manual → hourly with AI-assisted induction.

Three takeaways:

In vertical, low-tolerance, well-defined tasks, small models can beat general LLMs if data, training paradigm, and objective align with business.

Engineering investment in harness, skill governance, evidence anchoring beats endless prompt engineering for reliability.

Self-evolution must keep human decision authority : AI proposes, legal confirms, system logs, feedback splits into rule fast track and model slow track.

Key results summary
Key results summary

Q&A Highlights

Harness origin? Built in-house around existing tooling, not based on open-source agent frameworks.

Number of skills for fund promotion? 20+ skills , essentially split by risk type; broader business totals hundreds but loaded on-demand per business/entity/condition.

When to vertical fine-tune? High volume, high difficulty, clear boundaries, rule-based evaluation, sufficient scene data. Used 32B model . Observed clear "seesaw": vertical gains reduce general capabilities; if both needed, separate models.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI AgentReinforcement LearningFinancial Compliance32B ModelDual-Track EvolutionEvidence AnchoringMultimodal ParsingSkill-Based Rules
DataFunTalk
Written by

DataFunTalk

Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.