From Copilot to Loop: AI-Driven Legacy Refactoring via Data Cut Planes

Tencent's ad platform evolved from human-led refactoring with AI Copilot to Loop Engineering—a closed-loop system using data cut planes for observable verification and four specialized agents to automate refactoring of a 300k-line service, cutting cycle from months to weeks and manual diffs from 9 to near zero.

Tencent Advertising Technology
Tencent Advertising Technology
Tencent Advertising Technology
From Copilot to Loop: AI-Driven Legacy Refactoring via Data Cut Planes

Background: The dpcreative Monolith

The core creative service dpcreative has evolved over 10+ years to ~300k lines. It handles two key abstractions: DC (Dynamic Creative, advertiser-submitted) and TID (Traffic-facing ID, expanded per traffic spec). On average 1 DC expands to ~5 TIDs via a complex path involving spec validation, component composition, element matching, admission rules, and landing-page checks.

The monolith contains tightly coupled modules: public logic, DC logic, TID expansion, component logic, and element logic. Any cross-module architectural split carries high compensation risk due to coupling. Traditional "human-led split + dual-write reconciliation" hit three bottlenecks: understanding bottleneck (logic scattered across 300k lines), verification bottleneck (dual-write diffs require manual scheduling, making humans a serial throughput point), and collaboration bottleneck (high-priority demands pull engineers away, causing diff accumulation and code drift).

Evolution Overview: From L0 to L3

The team frames the journey as a paradigm shift across four levels (L0–L3) on axes of automation degree and verification observability. L3—the target—defines a Loop Engineering closure logic: diff = 0 → end / auto-promote (no human). diff > 0 && round ≤ 10 → continue Loop (Coding ↔ Reconciliation iterate). diff > 0 && round > 10 → human investigates (avoid token burn).

The project timeline shows four stages: 1a Element (human + Copilot → dpelement, done), 1b Component (human + Copilot → dpcomponent, reconciling), 2a TID v1 (entry split + import public lib → dpexpansion, done but risky), 2b TID v2 (Loop Engineering → dpexpansionv3, in testing).

Phase 1: Human-Led Domain Split with AI Copilot

Element Refactoring (First AI Copilot Scale Validation)

Challenge: ~2–3 engineers; 8 services, 60+ modification points; underlying storage incompatible compression, requiring data consistency. Steps: (1) Abstract element interface to shield table structure; 60+ mods done by AI Copilot, 3 weeks → 1 week. (2) Bidirectional reconciliation; reconciliation program AI-generated, 10 days → 3 days; found 4 diffs, manually fixed. Investment: ~2–3 people × (1 month dev + 2 months reconciliation), now fully rolled out.

Component Split: Structural Pain of Long Reconciliation Cycles

50+ cross-team mods; started 2025.06, paused for DB capacity work, restarted 2026.01. 15 legacy diffs blocked by serial human scheduling. Result: dpcreative reduced ~8.3k lines (~−2%); 8+ modules unified element/component write logic; but cycle stretched from 2025.06 to 2026.06 still unfinished. Key observation: reconciliation cycle length and developer pull-away are inevitable; fixing diffs by re-scheduling only worsens the cycle problem.

Phase 2a: The Public-Library Import Lesson

Ideal target: move TID expansion to dpexpansion calling dpcreative via RPC. Compromise: first version only split entry, reused public logic via import (not domain split). This caused two production incidents:

Panic : public logic changed, dpexpansion not adapted → ~2 engineers × 11 versions × 3 weeks.

Config fault : released dpcreative but forgot dpexpansion → production fault.

Lesson: public-library import is "fake split" —code coupling equals release coupling, cannot support high-risk refactoring. Gain: −3k lines (−0.9%); TID entry migrated; 1 person × 1 week dev + 1 month canary, now fully rolled out.

Paradigm Shift: Loop Engineering

From Vibe Coding to Closed-Loop Engineering

Fundamental difference: where the human sits . In Vibe Coding, human is execution center and verification blocker every iteration. In Loop Engineering, human becomes loop establisher; machines run 7×24 iterations with a 10-round circuit breaker to prevent token burn.

The loop comprises two sub-loops: Development Loop (compress dev + diff-fix cycle) and Testing Loop (replace manual case-by-case verification with cut-plane diffs). Full mechanism: human fixes "cut plane + acceptance criteria + structured input"; machine enters auto-iteration zone; diff=0 immediate end/auto-promote; diff>0 & round≤10 continue; diff>0 & round>10 human investigates.

Data Cut Planes: The Foundation of Loop

Whether Loop Engineering can close depends on one thing: can the legacy system be sliced into a set of "observable, comparable" intermediate states? A well-defined cut plane lets machines self-prove correctness and converge 7×24; a poor one reduces Loop to random token-burning search.

Methodology (3 steps + 5 quality gates): ① Speak human language → ② Find landing points → ③ Define schema+key; then apply five criteria to produce diff-able cut planes.

Definition & Three Mandatory Properties: A Data Cut Plane is a named observation point on the execution path specifying what business data to capture. Every cut plane on both old and new paths must be:

Recordable : old path can fully capture the data at that point as baseline.

Replayable : new path can deterministically re-run to that point under identical input .

Comparable : data has stable serialization format + alignment key for field-by-field diff.

A cut plane point Pᵢ = (locᵢ, schemaᵢ, keyᵢ) where locᵢ is code location (e.g., after function return, after DB write), schemaᵢ is the field set (by business semantics, not full memory dump), keyᵢ is alignment key (e.g., ad_id + creative_id). Reconciliation: for each Pᵢ, pair records by keyᵢ, compare field-by-field; if all pairs match, diffᵢ=0. Full-chain pass = all diffᵢ=0.

Five Criteria for a Good Cut Plane

Semantic completeness : sits at a "one thing done" business boundary (one DB write, one RPC return, one state change), explainable in human language.

Data self-contained : cut-plane data stands alone, no dependency on volatile runtime context (connections, sessions, memory pointers).

Stable alignment key : business primary key exists to pair records; unordered collections (lists) must be canonicalized (sorted) before pairing.

Determinism : strip non-deterministic fields—timestamps, auto-increment IDs, random numbers, concurrent write order—otherwise every replay produces false diffs.

Decoupling : adjacent cut planes should be low-coupled so diff at Pᵢ doesn't mechanically cascade to Pᵢ₊₁; ideally each cut plane is independently decidable.

First four decide "can the cut plane be compared"; fifth decides "is it easy to fix".

How to Define Cut Planes: From Human Language to Executable Spec

Cut-plane definition is the human-led step where cognitive contribution is densest. Three sub-steps:

Speak human language : without looking at code, list in business terms "what this service does". TID expansion: cross-product component combos → spec adaptation → write creative lib → write element lib → attach associations/async tasks → produce TID.

Find landing points : map each "one thing" to an observable code location. Three best location types: RPC boundaries (cross-service call inputs/outputs), DB write points (every write to adcreative, element, etc.), State/async change points (UpsertMpExtra, add74JobIfNeed, OnAdCreativeAdd side-effects).

Schema & key : for each landing point, decide fields and alignment key; Planning Agent outputs structured description; Coding Agent instruments accordingly.

The real TID expansion yielded 11 cut planes; cut planes 2–10 map almost one-to-one to a DB write or side-effect call.

Single Cut Plane Anatomy: Record–Replay–Compare

Old path captures schemaᵢ at cut point, normalizes (de-noise + sort), stores as baseline. New path captures same, normalizes identically, then diffs field-by-field against baseline by keyᵢ. Normalization is the key gatekeeper that keeps "false diffs" out.

Granularity: Inter-Cut-Plane "Step Count" Determines Success

Beyond where to place cut planes, how dense matters. More unobserved steps between adjacent cut planes → larger diff attribution scope → harder for AI to fix. Experiment E1 gave a tight counter-example:

2 cut planes, 9-step gap : 9 unobserved RPCs/writes between cut plane 1 and 2; diff appears at cut plane 2, root cause could be anywhere in those 9 steps—AI blind-guessed 50 times, still FAIL.

11 cut planes, 1-step gap : promoted those 9 steps into individual cut planes; each diff's attribution scope shrank to one step; AI aligned each, full-chain PASS.

E2 further showed: if cut-plane refinement is constrained, strong context (field mapping + taboo rules) can compensate for coarse granularity (see §7 Figure 7-1). Rule of thumb: "cut-plane spacing ↔ context strength" are interchangeable levers—first minimize spacing, then thicken context where spacing cannot shrink.

Difference from Unit / Snapshot Tests

Cut planes resemble snapshot tests but differ fundamentally:

Assertion target : Snapshot test = function output; Cut plane = business-process data state.

Golden source : Snapshot test = developer hand-written; Cut plane = recorded from old production path.

Focus : Snapshot test = is result correct?; Cut plane = is every step equivalent to old path?

Thus cut planes naturally fit behavioral-equivalence refactoring —goal is not "result matches" but "new path matches old path at every step".

Multi-Agent Collaboration Architecture

Why Strict Role Separation?

When a single agent merges Coding and Reconciliation, reconciliation code tampering appears—modifying verification logic to make diffs pass, destroying acceptance credibility.

Four-Agent Architecture

Decision Agent (Strong model): Orchestrate scheduling, Loop termination judgment. Must not write business code directly.

Planning Agent (Strong model): Cut-plane identification, structured requirements output. Must not execute reconciliation.

Coding Agent (Strong model): Refactoring development & diff fixing. Must not modify reconciliation program.

Reconciliation Agent (GPT-5.5): Independent replay, cut-plane-level diff. Must not modify business code.

Design principle: trade orchestration complexity for acceptance credibility —Reconciliation and Coding agents mutually constrain (see Figure 6-1). Decision Agent routes: diff=0 end/auto-promote; diff>0 & round≤10 continue fix; round>10 human investigates.

Controlled Experiments & Lessons (Core Chapter)

E1: Cut-Plane Granularity — 2 vs 11 Cut Planes

Motivation: too many unobserved steps between cut planes → AI cannot locate diff root cause. Design: same task, two setups. Result: 2 cut planes (9-step gap) → 50 attempts / 1 day unresolved → FAIL. 11 cut planes (1-step gap) → incremental PASS → full-chain PASS. Conclusion: shrink cut-plane spacing → pass each → full-chain pass. Counter-example: 9-step gap makes intermediate state invisible, AI can only blind-guess.

E2: Cut-Plane Granularity × Context Strength (2D Trade-off)

Motivation: E1 suggests "fine cut planes lower input demand"—must we use 11? Can coarse cut planes + strong context achieve parity? Design: keep 2 cut planes, upgrade Planning Agent output to: (1) function responsibility matrix (sequence diagram + responsibility map), (2) field mapping graph (source→target field + code location), (3) engineering taboo rules. Result: Control (weak ~800-word TAPD) → 50 attempts FAIL, high human intervention. Experiment (structured semantic model) → 2 attempts PASS, human intervention ≈0. Conclusion (2D solvability formula): Solvability ≈ f(cut-plane-granularity⁻¹, context-strength). Two equivalent paths: (1) coarse cut planes + strong context (E2 path); (2) fine cut planes + weak context (E1 path).

E3: Build Strategy — Greenfield vs. Move-Then-Delete

Object: TID expansion Prepare path. Comparison:

Core business code : Greenfield 1796 lines / 17 files; Move-Then-Delete 2884 lines / 16 files.

External call code : Greenfield separate 500 lines; Move-Then-Delete mixed 464 lines.

Main function length : Greenfield 42 lines; Move-Then-Delete 82 lines.

Review difficulty : Greenfield High; Move-Then-Delete Low.

Reconciliation failure diagnosis : Greenfield Hard; Move-Then-Delete Easy.

Time to usable : Greenfield Slow; Move-Then-Delete Fast.

Quality ceiling : Greenfield High; Move-Then-Delete Medium.

Conclusion: For reconciliation-loop scenarios, prefer move-then-delete to enter Loop fast; long-term governance target remains greenfield architectural quality.

E4: Architecture Optimization Skill Evolution (v1→v2→v3)

Motivation: upgrade "luck-based code reading" to reproducible, auditable architecture analysis capability.

v1 : Free LLM analysis → 7 optimizations found; numbers estimated, no line numbers.

v2 : Scripted metrics + LLM interpretation → 9 optimizations (+2); objective metrics + line-number traceability.

v3 : v2 + rewrite strategy option → 12 optimizations (+3); API freeze / internals free to change.

Upgrade drivers: v1→v2: same repo two runs gave inconsistent results → script the mechanically repeatable parts. v2→v3: ignored "downstream has reconciliation safety net" → overly conservative → introduced rewrite option.

Effectiveness Evaluation: TID Expansion Second Attempt

Timeline

Infra : 2026.04.30–06.03 — record/replay pipeline + agent workflow.

Refactor : 2026.06.04–06.16 — cut plane 1 diff zeroed (06.10) → cut plane 2 submitted for test (06.16).

vs. Traditional Approach

Dev + reconciliation cycle : Traditional 1–2 months; Agent Workflow 2 weeks .

Diff handling : Traditional all manual scheduling; Agent Workflow 35 diffs: 26 auto, 9 human (after structured TAPD, human ≈0).

Dev-phase headcount : Traditional ongoing multiple; Agent Workflow 1 person (infra ~2–3 people one-time).

Future expectation (similar projects) : ≤1 week / 0.5 person if refactor has no behavioral change.

Role : Traditional code worker; Agent Workflow architect + rule definer.

Generalizable AI Architecture Refactoring Workflow

Four reusable steps (Figure 9-1):

Human does architecture analysis.

Human + Planning Agent co-complete cut-plane definition + structured requirements (cut-plane loc/schema/key + responsibility matrix / field mapping / taboo rules are the same thing—describing service semantics—hence merged).

Infra completes record/replay pipeline.

Four-agent loop — diff=0 end/auto-promote; only >10 rounds triggers human investigation .

Transferability layering (Table 9-1): Paradigm, Governance, Infra, Process, Levers are domain-agnostic and directly reusable; only Domain layer (cut-plane definition + reconciliation semantics) requires domain experts + Planning Agent. Main incremental cost for cross-domain adoption is the two domain-layer rows; paradigm/governance/infra/process need no rebuild. This article uses ad-tech creative pipeline as verification scenario, not method boundary.

Conclusion: Five Principles

The dpcreative refactoring shows: AI-era large-system refactoring essence is not "let AI write more code" but refactor the R&D operating system .

Acceptance boundary first — data cut planes turn "understand business" into diff-able engineering objects.

Closed-loop beats open-loop — diff=0 end/auto-promote; only diff>0 & >10 rounds human investigates, avoiding token burn.

Governance over brute force — four-agent role separation is the institutional guarantee of acceptance credibility.

Input determines output — cut-plane granularity × context strength form a 2D lever (E1+E2).

Infra once, reuse forever — record/replay + agent workflow are amortizable assets.

From 300k-line monolith to domain-oriented service cluster; from ~2–3 people × months to 1 person × 2 weeks — this is the paradigm migration of engineers from code workers to architects and rule definers .

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI-assisted developmentarchitecture evolutionMulti-Agent Collaborationlegacy refactoringLoop Engineeringbehavioral equivalencedata cut planeTencent advertising platform
Tencent Advertising Technology
Written by

Tencent Advertising Technology

Official hub of Tencent Advertising Technology, sharing the team's latest cutting-edge achievements and advertising technology applications.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.