ART: Two-Stage Makeup Transfer Anchors Supervision to Real Images, Beats Pseudo-Target Ceiling

vivo BlueImage Lab introduces ART, a two-stage makeup transfer framework that first learns from synthetic pseudo-targets then refines using real reference images as supervision, achieving state-of-the-art fidelity on complex makeup like glitter and face painting while preserving identity, and releases the first 2K makeup dataset MF2K.

vivo Internet Technology
vivo Internet Technology
vivo Internet Technology
ART: Two-Stage Makeup Transfer Anchors Supervision to Real Images, Beats Pseudo-Target Ceiling

The article presents ART (Accepted at ECCV 2026), a makeup transfer framework from vivo BlueImage Lab that addresses the fundamental lack of paired training data in makeup transfer. The core problem: ideal supervision would require paired images of the same person with and without makeup under identical conditions, which is physically infeasible at scale.

Two existing workarounds are analyzed: (1) GAN-based methods using weak priors like cycle consistency, which capture color and ambience but fail on spatially precise elements (glitter, stickers, face painting) because the supervision signal never contains such information; (2) Diffusion-era pseudo-target generation, where a large image-editing model synthesizes a made-up result from a source and reference, then a student model is trained to mimic it. This converts the task to supervised learning but introduces a structural flaw: every error in the pseudo-target — missing glitter, altered facial structure — is inherited by the student, creating a "pseudo-target ceiling" that caps performance.

ART's solution redefines the role of pseudo-targets: they are only an initialization. The method uses two stages:

Stage I – Pseudo-target Initialization: Train a transfer model on (source, reference, pseudo-target) triplets to learn basic spatial distribution, color mapping, and sticker placement. Simultaneously train an auxiliary makeup removal model that reconstructs a bare face from a made-up image — this model becomes critical in Stage II.

Stage II – Reality-Anchored Refinement: The transfer model's output is treated as a differentiable "makeup carrier" (a mathematically differentiable makeup layer). This layer is removed from the source face and composited onto the reference's bare-face version (produced by the Stage I removal model). The reconstruction is compared to the real reference image; any discrepancy — missing glitter, misplaced patterns, boundary artifacts — produces a loss that backpropagates through the differentiable carrier to the transfer model. Thus supervision comes directly from the real image, not from imitation of a synthetic target.

A key hyperparameter is the noise level applied to the pseudo-target before refinement. Experiments show 0.6 is optimal: enough freedom to correct details without destabilizing identity and background.

The paper also introduces MF2K (MakeupFaces2K) , the first 2K-resolution real-world makeup dataset. From ~100K high-res portraits plus FFHQ originals, after face alignment, CLIP deduplication, quality/pose filtering, VLM annotation, and manual review, MF2K contains 8,573 images at 2048×2048 across four categories: bare (3,139), light makeup (2,063), heavy makeup (1,798), artistic makeup (1,573). High resolution preserves eyelashes, glitter particles, lip texture — details that vanish at lower resolutions.

Experiments on MT, LADN, MT-Wild, and MF2K (400 non-overlapping identity pairs; MF2K test set uses only artistic makeup) compare against three generations: GAN (PSGAN, EleGANt), specialized diffusion (StableMakeup, SHMT, MAD), and commercial models (Nano Banana Pro, GPT Image 1.5). Metrics: makeup similarity via two vision models (MSimG, MSimQ), identity consistency (ID), background change (L2-M), overall quality (FID). ART achieves top MSimG on all four test sets (9.22 on MF2K artistic vs. 8.43 next best) and lowest L2-M everywhere. It does not win every metric (MSimQ and FID have other leaders) but uniquely pushes makeup fidelity and background stability together while keeping identity and quality in the top tier. User study (21 participants, 864 blind ratings) ranks ART first on makeup similarity (3.67), identity consistency (4.03), and image quality (3.83) out of 5.

Two direct verifications confirm the pseudo-target ceiling is broken: (1) On the same training samples, ART corrects head-pose shifts, wrong patterns, and semantic hallucinations (e.g., pseudo-target generated non-existent letters on forehead) present in its own training targets, producing outputs superior to the pseudo-targets. (2) Replacing the strong pseudo-target generator (Nano Banana Pro) with a weaker one (StableMakeup, MSimG 6.73) still yields an ART variant scoring 8.56; switching back to the strong generator reaches 9.22. Pseudo-target quality affects the starting point, not the ceiling. Ablation shows Stage I alone gives MSimG 8.82, ID 0.57, L2-M 7.68; adding Stage II improves to 9.22, 0.74, 4.31 — gains come from real-image supervision, not data or initialization. Retraining StableMakeup on MF2K only reaches 6.58, ruling out data advantage.

ART is the first makeup transfer framework supporting native 2048×2048 output. Training at 512, 1024, 2048 shows progressive sharpening of sticker edges, glitter granules, and eyelashes; a wavelet-transform loss at high resolution further strengthens high-frequency makeup structure learning, confirming 2K output is trained, not upscaled.

Limitation: ART does not explicitly model lighting and material. When source and reference lighting differ strongly, strong highlights from the reference may be transferred into mismatched illumination, creating unrealistic local effects (Figure 9).

The authors argue this work signals a shift in portrait enhancement from generic filters to detail fidelity, and that the two-stage paradigm — synthetic initialization followed by real-image anchored refinement — is transferable to any image editing task starved of paired data (hairstyle change, wrinkle removal, age progression). The key insight: pseudo-targets are a stepping stone, not the destination; a differentiable path from model output to real-image consistency turns "disagreement with reality" into a loss function.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

computer visionidentity preservationtwo-stage trainingECCV 2026high-fidelity generationmakeup transferMF2K datasetpseudo-target ceiling
vivo Internet Technology
Written by

vivo Internet Technology

Sharing practical vivo Internet technology insights and salon events, plus the latest industry news and hot conferences.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.