From Top‑Conference Papers to Real‑World AI Imaging: Meitu MT Lab’s Quest to Make AI Truly Useful

Over the past two years Meitu’s MT Lab has published eleven top‑conference papers on video editing, portrait manipulation, visual understanding and more, and has transformed these research breakthroughs—such as MiVE, CFT, BridgeRemoval, WearWow, FlowSeg and All‑in‑One Slider—into product features that let AI grasp creators’ intent and deliver precise, controllable image and video edits.

Machine Heart
Machine Heart
Machine Heart
From Top‑Conference Papers to Real‑World AI Imaging: Meitu MT Lab’s Quest to Make AI Truly Useful

While generating images and videos has become easy, the real difficulty lies in making AI understand and follow creators’ true intent, such as editing text without breaking layout, removing redundant elements seamlessly, preserving consistency across frames, and enabling realistic virtual try‑on.

Prompt engineering often adds long constraint phrases ("keep consistency, avoid distortion, do not change"), but the ideal is for AI to execute the request directly without extensive prompting.

Long‑Term Research at Meitu MT Lab

Since 2024, Meitu’s MT Lab has published eleven papers at top venues (ICLR, CVPR, ICML, ECCV, ACM MM) covering video editing, portrait editing, intelligent interaction, and visual understanding. These works are continuously transferred into product capabilities.

MiVE – Reference‑Guided Video Editing

Presented at ICML 2026, MiVE introduces a multiscale vision‑language framework with three components: hierarchical visual‑language context extraction, reference‑frame‑aware latent encoding, and a unified self‑attention backbone. It uses multiscale VLM conditioning to overcome long‑standing structural gaps in instruction understanding and spatial precision.

MiVE preserves motion structure in large‑movement, heavy‑occlusion, and cross‑frame style‑consistency scenarios, and shows stronger stability in identity, detail sharpness, and color consistency compared with prior models.

CFT – Consistent Feature Transport for Relighting

At ECCV 2026, CFT reframes portrait relighting as a "consistent feature transport" problem. Built on the Rectified Flow framework, it explicitly learns lighting transformations between source and target distributions using a large‑scale relighting dataset, achieving high fidelity in complex lighting scenes.

CFT provides fine‑grained control over light direction, angle, and intensity, distinguishing morning from sunset light and delivering stable, realistic illumination even in challenging scenes.

BridgeRemoval – Video Target Removal via VP‑SDE Bridge

Introduced at ICML 2026, BridgeRemoval treats video object removal as a video‑to‑video translation task using a variance‑preserving stochastic differential equation (VP‑SDE) bridge. It selectively transforms masked regions while preserving unmasked background fidelity.

The method reduces processing time by about 80 % and supports unlimited video length, outperforming mainstream models that are limited to 30‑second clips.

WearWow – 2K Multi‑Garment Virtual Try‑On

At ECCV 2026, WearWow tackles 2K‑resolution multi‑garment try‑on with Adaptive 2D Token Packing (ATP) for computational efficiency and Multi‑dimensional Try‑on Reward (MTR) for direction adjustment. It preserves material texture (wool, denim, knit) and handles complex poses and occlusions.

WearWow achieves realistic dressing rather than simple garment replacement, maintaining high‑resolution texture fidelity even with multiple garments.

All‑in‑One Slider – Unified Multi‑Attribute Control

Presented at CVPR 2026, the All‑in‑One Slider expands the one‑for‑one attribute‑slider paradigm to a unified framework that continuously controls age, smile, makeup, hairstyle, etc., supports zero‑shot generalization to unseen attributes, and preserves identity and background consistency.

FlowSeg – Bidirectional Semantic Flow for Conditional Segmentation

At ICML 2026, FlowSeg redefines the role of language in LLM‑conditioned segmentation. Language guides feature updates at every decoding layer, ensuring the mask stays aligned with the textual description throughout generation.

FlowSeg achieves accurate mask‑language alignment even in highly complex scenes, reducing the “segment‑right‑but‑select‑wrong” failure mode of conventional propose‑then‑select pipelines.

PE‑Field – 3D Positional Encoding for Diffusion

In ICLR 2026, the Positional Encoding Field (PE‑Field) extends 2D positional encodings to 3D by incorporating depth information, enabling finer spatial control from patches down to sub‑patch regions.

PE‑Field improves single‑image novel‑view synthesis, object rotation, and spatial adjustments, laying a solid foundation for multi‑view visual content creation.

ControlHair – Physically‑Based Dynamic Hair Rendering

At ECCV 2026, ControlHair couples a physical simulator with a video diffusion model via a cascade architecture, decoupling physics from image generation to achieve controllable, realistic dynamic hair that follows physical laws.

More than 50 hair‑related AI features (e.g., “hair flutter”, “glowing strands”) have been launched, demonstrating the framework’s versatility.

ACaM – Atomic Camera Movement Understanding

In ACM MM 2026, the MT Lab defined natural‑language camera‑movement understanding as an independent research problem, identified five failure modes of VLMs, and released the ACaM benchmark covering 17 movement types across real and synthetic videos.

Fine‑tuned VLMs on a balanced 27k‑sample dataset show significant gains in camera‑movement comprehension.

SI‑Edit – Collaborative Text‑and‑Sketch Image Editing

At ACM MM 2026, SI‑Edit fuses textual instructions with sketch constraints using learnable task triggers and shared positional encodings, trained on ~6.5k quadruplets (source, target, sketch, command) to enable precise pixel‑level edits such as scaling flowers, bending branches, or shaping flames.

These advances illustrate how Meitu’s MT Lab moves AI imaging from mere generation toward true creation, continuously raising model capabilities while delivering products that users not only can use but love.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

researchvideo editinggenerative AImultimodal diffusionvisual understandingAI imagingportrait editing
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.