Tencent's 32B Financial Compliance Agent: Skill Rules & Dual-Track Self-Evolution
Tencent's financial compliance agent uses a 32B model, skill-based rules, and dual-track self-evolution to achieve 95%+ recall, 80%+ precision, 40-second inference, and hourly rule updates, replacing manual review that took days.
Why Manual Review Fails
Financial marketing compliance review is not just checking copy. Two regulatory cases illustrate the problem: process failures (posters launched without proper compliance checks, missing risk disclosures and personnel qualifications) and selective disclosure (showing only partial strong performance). Both point to the same issue: a single compliance decision must verify content, data, entity, and process simultaneously, forming a complete evidence chain .
Review objects have expanded from text-only (customer service messages, app pop-ups) to posters, banners, long product pages, short videos, and live streams. The traditional linear workflow—receive materials, queue dispatch, human judgment, spot-check/review—means material volume growth directly demands proportional headcount growth . This leads to insufficient coverage, review cycles measured in days or weeks, and inconsistent judgments across reviewers. New regulations also create a window where old practices persist.
The system was redefined with four goals: see everything, judge fast, judge accurately, keep up —unifying multi-material intake, pre-publish review, stable re-review, and rapid new-rule effectiveness into system capabilities.
Turning Regulations into Agent-Executable Rules: From Long-Image Parsing to Evidence Chain
Online Pipeline Architecture
The online chain: incoming materials → multimodal parsing + entity/scene identification → skill rule routing, risk reasoning, evidence anchoring, reverse verification → output risk tags, levels, violation evidence, cited clauses, and modification suggestions. Supporting capabilities include perception parsing, tool calling, rules & strategies, compliance reasoning, knowledge & memory. Offline handles skill governance, rule self-evolution, and model data flywheel.
Perception Parsing: Long-Image Parsing Challenges & Solution
Fund promotion long pages contain product info, performance charts, manager bios, and bottom risk disclosures—requiring content completeness while preserving visual semantics and spatial relationships. Pure OCR loses charts and positional relations; general multimodal models miss bottom fine print, hallucinate text, or fail to parse complex charts fully.
Engineering solution: three-stage processing
Semantic slicing by visual boundaries, background changes, and OCR text boxes; adjacent slices retain overlap.
Feed each slice to multimodal model to extract text, charts, elements, and positional relations.
Sort by original image coordinates, deduplicate overlaps, and semantically stitch.
Results: recall improved from 90.15% to 97.59% , F1 from 94.73% to 98.67% , miss rate dropped from 9.85% to 2.41% , average latency increased from 16s to 21s.
Rule Sources & Structuring
Rule base comprises three layers: basic laws, financial industry special rules, content & AI governance , continuously ingesting regulatory bulletins, penalty cases, and industry self-discipline rules. Each rule carries regulation/file number, original text, applicable entity & scenario, effective status, version, and case basis—ensuring traceability and updatability.
Why Raw Regulations Can't Be Used Directly
Feeding raw regulations to an agent hits five execution breakpoints : content too long/scattered, abstract execution boundaries, unclear applicability scope, missing evidence & disposition structure, and version drift. Corresponding engineering transformation also has five steps:
Map by entity, product, channel, time.
Atomicize clauses into entity, action, object, condition, exception, measurable thresholds.
Build applicability matrix: entity × product × channel × audience × action.
Add decision conditions and human routing.
Validate with positive, negative, and boundary examples before versioned release.
In fund promotion, tag system handles routing, structured rules handle adjudication , covering promotional claims, historical performance, index data, risk disclosure, and licensed business risks.
Skill-Based Governance
As rules grow, stuffing all into the prompt causes linear token cost growth and attention dilution. Adopted two-level progressive disclosure : core risk rule definitions stay in prompt for high-recall initial screening; on risk hit, load the corresponding skill's boundaries, exceptions, and historical cases. Each inference sees only the rule subset relevant to the current material.
Tool Calling for Auditable Evidence
Compliance conclusions rely on dynamic facts outside the material. Example: "fixed income plus" promotion needs actual holdings data not in the image. System performs entity disambiguation , plans data sources, query fields, and valid time points, calls external APIs via MCP parameterized invocation; results validated for type, unit, range, source → form evidence snapshots with timestamp, tool version, raw response. Timeouts, missing data, or source conflicts escalate to human review.
Evidence Anchoring: The Deliverable Is Evidence, Not Just Judgment
System maps conclusions back to original image coordinates: text via character boxes, visual objects via open-vocabulary object boxes, irregular regions via pixel contours—using EasyOCR/CRAFT, Grounding DINO, SAM. Localization failures or evidence conflicts also fall back to human review.
Why Train a Small Model: Financial Compliance Needs Not "General Capability"
Problems with General LLMs
Vertical risk identification accuracy/recall only ~30-40% .
Large models: single inference 128 seconds on 16 GPUs .
General training objectives don't align with compliance's "high recall first" requirement.
Average ~15k token long prompts easily lose focus.
Team trained a task-specific model targeting
high precision/recall, low latency, long-context localization
.
Three-Stage Training
Data synthesis (dual-path):
Reverse synthesis: from human-confirmed gold-risk samples → correct to risk-free text → generate new samples per risk label.
Forward synthesis: unlabeled materials → OCR, long-image slicing, multi-model labeling merge → human verification.
SFT : teaches review requirements and structured output, kept lightweight to avoid over-training.
Reward models & RL : based on full compliance quadruples and negative examples (missed, false, hallucinated) construct pair preference data ; train model-based and rule-based reward models.
Three Key Decisions in SFT/RL Combination
Data coverage: native/synthetic, risky/safe, long-tail categories—avoid head-risk dominance.
Keep SFT light to leave exploration space for RL.
Base model selection: not just Pass@1, but Pass@8 to observe potential ceiling .
Business-Adjusted Objective: F2 Score
Standard F1 weights precision and recall equally. Compliance fears misses more, so adopted F2, making recall weight 4× precision .
Dual-Track Self-Evolution: Rule Changed Today, Capability Grows Tomorrow
Rule Self-Evolution (Fast Track)
"Self-evolution" does not mean the agent rewrites rules and deploys directly. Inputs: misjudgments/misses from legal review, proactively scanned regulatory signals, rule conflict detection. AI inductively identifies uncovered rules, missing exceptions, fuzzy boundaries and generates suggestions; legal confirms, corrects, or rejects. Confirmed rules enter rule library, next inference dynamically loads them—no model retraining needed. Updates achievable in hours .
Dual-Track Evolution
Rule track : solves "today's problem fixed today"—legal feedback directly becomes skill updates for rapid correction.
Model track : distills experience into capability—confirmed/corrected/rejected labels become RL signals for periodic GRPO retraining.
Both tracks share the same real feedback source but handle short-cycle rule changes and deep model capability separately.
Results & Conclusions
Coverage: from sampling to full coverage .
Cycle: 3-5 days → minutes .
Custom 32B model: recall 95%+, precision 80%+ .
Inference: 128s (671B) → 40s (32B) , ~69% reduction; deployment compute cost -75% .
Rule updates: weekly+ manual → hourly with AI-assisted induction.
Three takeaways:
In vertical, low-tolerance, well-defined tasks, small models can beat general LLMs if data, training paradigm, and objective align with business.
Engineering investment in harness, skill governance, evidence anchoring beats endless prompt engineering for reliability.
Self-evolution must keep human decision authority : AI proposes, legal confirms, system logs, feedback splits into rule fast track and model slow track.
Q&A Highlights
Harness origin? Built in-house around existing tooling, not based on open-source agent frameworks.
Number of skills for fund promotion? 20+ skills , essentially split by risk type; broader business totals hundreds but loaded on-demand per business/entity/condition.
When to vertical fine-tune? High volume, high difficulty, clear boundaries, rule-based evaluation, sufficient scene data. Used 32B model . Observed clear "seesaw": vertical gains reduce general capabilities; if both needed, separate models.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunTalk
Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
