SmartPhotoCrafter: Think-First Photo Enhancement Unifies Restoration & Retouching

SmartPhotoCrafter introduces a unified reasoning-to-generation framework for automatic photo enhancement: an Image Critic analyzes quality via chain-of-thought reasoning, then a Photographic Artist performs high-fidelity restoration and aesthetic retouching, unifying dehazing, deblurring, and color grading without explicit user instructions, trained via a three-stage strategy with multi-level rewards.

vivo Internet Technology
vivo Internet Technology
vivo Internet Technology
SmartPhotoCrafter: Think-First Photo Enhancement Unifies Restoration & Retouching

SmartPhotoCrafter, developed by vivo BlueImage Lab, presents a novel unified understanding–generation–optimization framework for automatic photographic image enhancement. The method models image optimization as a "reasoning-to-generation" pipeline: first analyze, then decide, finally enhance. It addresses two key challenges in existing approaches: the lack of explicit aesthetic defect perception and structured reasoning, and the disconnect between restoration (e.g., dehazing, deblurring) and retouching (e.g., color grading, style optimization).

Research Background: Let the Model Think First, Then Edit

Current automatic photographic image editing has evolved from instruction-driven local edits to intelligent aesthetic enhancement without explicit instructions. However, models still struggle to diagnose image quality issues and formulate optimization plans like a photographer. Most methods rely on end-to-end conditional generation or instruction-based local adjustments, requiring users to specify edits such as "make the sky bluer" or "remove haze." The real difficulty lies in discovering problems and devising reasonable enhancement strategies.

Core Technology: SmartPhotoCrafter Framework

2.1 Overall Framework

The framework consists of two highly collaborative modules:

Image Critic (图像评论家): The "aesthetic brain" responsible for reasoning and judgment. Given an input photo, it performs comprehensive quality analysis via Chain-of-Thought (CoT), outputting three components: a detailed quality analysis reasoning trace, a set of structured editing suggestions, and a 0–100 image quality score. It is built on the vision-language model Qwen2.5-VL-7B, supervised fine-tuned on image quality assessment (IQA) and photographic editing data to acquire strong photographic aesthetic understanding.

Photographic Artist (摄影艺术家): Executes the image optimization generation. Based on Qwen-Image-Edit-2509, it uses a Diffusion Transformer (DiT) architecture trained with Flow Matching. The Artist receives two inputs: the original image and the reasoning latent representation from the Critic. This tight coupling allows the Artist to understand what problems exist and which direction to optimize, generating enhanced results without relying on explicit textual instructions.

2.2 Three-Stage Training: Progressive Mastery

Stage I – Foundation Pre-training: Image Critic and Photographic Artist are trained independently. The Critic learns to "look at images and give aesthetic judgments," while the Artist learns basic "retouching capabilities" on restoration and retouching data pairs.

Stage II – Reasoning-Conditioned Adaptation: The core goal is to teach the Photographic Artist to understand the Image Critic's "thought process." Instead of using natural language as an intermediate bridge, the Artist directly consumes the Critic's reasoning latent representation, aligning the two modules at the representation level.

Stage III – Coordinated Reasoning-to-Generation RL: Moves from supervised learning to joint reinforcement optimization. The Image Critic is optimized with GRPO (Group Relative Policy Optimization) to produce more stable and instructive aesthetic judgments. The Photographic Artist employs DiffusionNFT, a method extending GRPO ideas to the continuous space of diffusion models.

2.3 Multi-Level Reward Mechanism for Collaborative Optimization

The third-stage reinforcement learning optimizes both modules end-to-end. The reward design emphasizes attribute-level parameter sensitivity , decomposing the hard-to-model overall image quality evaluation into fine-grained, interpretable attributes such as exposure, contrast, saturation, and color temperature. This attribute-decoupled design significantly enhances the optimization process's perception of subtle color-tone differences, enabling more stable and controllable responses to minor retouching variations. The complete reward architecture is illustrated in the figure below.

Multi-level reward mechanism diagram
Multi-level reward mechanism diagram

2.4 Multi-Task Dataset Construction

To support progressive learning from understanding to controllable generation to cross-module coordination, a phased multi-task dataset was built:

Image Critic training data: Sourced from IQA datasets and image restoration datasets. CoT annotations (quality analysis, editing suggestions, quality scores) were generated by Qwen2.5-VL-72B; IQA dataset scores replaced with original ground-truth values.

Photographic Artist training data: Includes restoration pairs (deblurring, dehazing, shadow removal) and retouching pairs (exposure, contrast, saturation, color temperature, depth of field), plus composite editing pairs created by randomly stacking multiple retouching operations on low-quality images.

Alignment stage data: In addition to the above, incorporates MIT-Adobe FiveK and a high-quality subset filtered from the AVA dataset. The AVA subset constructs "degraded–high-quality" pairs via synthetic perturbations (exposure shift, contrast compression, saturation distortion, color temperature changes) to strengthen optimization of low-quality images.

Dataset construction overview
Dataset construction overview

Performance: Image Enhancement Across Multiple Scenarios

SmartPhotoCrafter was tested on various image optimization scenarios, validating its automatic photographic enhancement capability.

Comparison with Other Methods on Automatic Photographic Enhancement

Compared to other methods, SmartPhotoCrafter achieves excellent restoration and retouching results while maintaining realism and content consistency. Qualitative results show superior dehazing, deblurring, and color grading.

Qualitative comparison with other methods
Qualitative comparison with other methods

Multi-Attribute Combined Editing

In multi-attribute instruction editing tasks, SmartPhotoCrafter jointly controls restoration, exposure, saturation, and other attributes, preserving structural and stylistic stability across diverse scenes.

Multi-attribute editing examples
Multi-attribute editing examples

Additional Results

Further visual results demonstrate consistent quality improvements across varied lighting and degradation conditions.

More enhancement results
More enhancement results

Conclusion

SmartPhotoCrafter unifies "image understanding + retouching" by connecting semantic understanding with image generation, achieving automatic photographic image optimization covering both restoration and retouching tasks. It delivers stable, natural-looking results across multiple enhancement scenarios. The current model focuses on low-level adjustments for restoration and color; it does not yet handle complex composition-level edits. Future work can extend toward "understanding composition and holistic optimization" for more comprehensive post-processing.

Paper: https://arxiv.org/abs/2604.19587 Project page: https://vivocameraresearch.github.io/smartphotocrafterweb Code repository:

https://github.com/vivoCameraResearch/SmartPhotoCrafter
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

chain-of-thoughtReinforcement Learningimage restorationdiffusion transformerautomatic photo enhancementcolor gradingreasoning-to-generationSmartPhotoCrafter
vivo Internet Technology
Written by

vivo Internet Technology

Sharing practical vivo Internet technology insights and salon events, plus the latest industry news and hot conferences.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.