BeautyGRPO: A New Reinforcement Learning Framework that Recreates Realistic Portraits
The CVPR 2026 paper introduces BeautyGRPO, a reinforcement‑learning framework that leverages the fine‑grained FRPref‑10K portrait‑retouching preference dataset and a novel Dynamic Path Guidance algorithm to simultaneously enhance skin texture, preserve identity features, and achieve superior aesthetic alignment, outperforming existing retouching models on objective metrics and user preference tests.
Background and Problem
High‑quality digital portrait retouching demands both precise removal of minute skin defects and preservation of native identity traits, creating a zero‑sum tension between "high fidelity" and "human aesthetic". Existing supervised fine‑tuning (SFT) models such as RetouchFormer and generic editors like NanoBanana over‑fit to pixel‑level references, leading to either residual flaws or overly smoothed, plastic‑like faces. Online reinforcement‑learning approaches (e.g., FlowGRPO) introduce stochastic exploration that causes cumulative random drift, severely degrading fidelity.
Core Contributions
1. FRPref‑10K Fine‑Grained Preference Dataset – The authors construct the first large‑scale dataset containing 10,000 high‑resolution portrait pairs annotated with five fine‑grained dimensions: skin smoothness, flaw removal, texture quality, clarity, and identity‑feature retention. Using visual‑large‑model (VLM) embeddings calibrated by human experts, they train a multi‑dimensional reward model capable of detecting subtle aesthetic differences such as micro‑texture and gloss variations.
2. Dynamic Path Guidance (DPG) – To reconcile aesthetic exploration with fidelity, DPG introduces an “anchor‑constraint” mechanism during sampling. At each diffusion step, a deterministic correction vector is computed by planning a trajectory toward a high‑quality reference anchor and blending it with the original stochastic differential equation (SDE) direction. The algorithm applies a time‑dependent weight decay:
Early (high‑noise) stage: Strong correction weight pulls the trajectory back onto the high‑fidelity manifold, stabilizing facial structure and lighting.
Late (detail) stage: The correction weight is gradually reduced, allowing controlled exploration to surpass the anchor’s aesthetic quality while staying within a safe fidelity boundary.
Experimental Evaluation
To avoid the “perception‑distortion” trade‑off of full‑reference metrics (e.g., PSNR), the authors adopt no‑reference aesthetic scores (NIMA, MUSIQ, MANIQA). Results show that BeautyGRPO consistently outperforms specialized retouching models and general‑purpose editors across all NR metrics. Identity preservation measured by ArcFace remains above 0.95, confirming that facial traits are not compromised.
In a double‑blind user study with 100 participants of varied ages and professional editing experience, BeautyGRPO achieved a 63.25% preference win rate, far ahead of the runner‑up (12.00%). The multi‑dimensional reward model’s scores correlate strongly with human ratings, demonstrating accurate aesthetic alignment.
Generalization
Applying BeautyGRPO to the generic Qwen‑Image‑Edit large model eliminates the latter’s typical “identity drift” and “over‑smoothing” issues during facial edits, evidencing strong plug‑and‑play generalization.
Conclusion
BeautyGRPO resolves the longstanding conflict between extreme aesthetic exploration and native fidelity, delivering portrait retouching that preserves personal traits while achieving natural‑looking skin texture. The work, accepted as a Highlight at CVPR 2026, showcases the potential of fine‑grained reward modeling and dynamic path guidance to advance computational photography and AIGC.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
vivo Internet Technology
Sharing practical vivo Internet technology insights and salon events, plus the latest industry news and hot conferences.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
