Unlocking Free‑View Video Virtual Try‑On with TryOnCrafter’s 4D Try‑On Proxy

TryOnCrafter introduces a camera‑controllable video virtual try‑on framework that builds a renderable 4D proxy to enable free‑view, 360° and bullet‑time effects while preserving structural stability, texture consistency, and realistic motion across arbitrary camera trajectories.

Machine Heart
Machine Heart
Machine Heart
Unlocking Free‑View Video Virtual Try‑On with TryOnCrafter’s 4D Try‑On Proxy

Research Background

Video Virtual Try‑On (VVT) aims to transfer a target garment onto a person in a video while keeping identity and motion coherent. Existing VVT methods are limited to the original camera trajectory of the source video, preventing users from freely rotating, zooming, or circling the subject to inspect the clothing from different angles.

Method Overview

TryOnCrafter defines a new task—Camera‑Controllable Video Virtual Try‑On (CaM‑VVT)—which requires generating videos that follow user‑specified camera paths while maintaining structural stability, texture continuity, natural motion, and consistent foreground‑background relationships.

The framework consists of two tightly coupled stages: (1) a renderable 4D try‑on proxy that provides an explicit geometric anchor, and (2) a video diffusion Transformer (Video DiT) that refines the proxy rendering into a photorealistic video adhering to the target camera trajectory.

Renderable 4D Try‑On Proxy Construction

The proxy is a time‑varying 4D representation built through:

Background reconstruction : recover background point cloud and camera parameters from the source video.

Human motion recovery : align the SMPL‑X body model to the source motion sequence.

2D try‑on prior distillation : distill high‑quality 2D try‑on results into a 3D Gaussian Splatting (3DGS) model, creating a clothed 3D avatar.

Scene alignment : anchor the SMPL‑X motion to the reconstructed background point cloud, ensuring consistent pose, scale, and perspective.

By unifying person, garment, background, and camera trajectory in a single 4D world space, each frame is constrained by explicit geometry rather than implicit pixel cues.

Video Diffusion Transformer Construction

After rendering the proxy along the target camera path, the initial video serves as a structural prior for a Video DiT based on the pre‑trained Wan2.1‑I2V‑14B model. The DiT refines texture, identity details, and occluded regions while strictly following the camera trajectory.

Three conditioning mechanisms are employed:

Rendered prior : the proxy rendering provides spatio‑temporal structure.

Cross‑view reference adapter (CRA) : extracts fine‑grained identity, garment, and background features from the source video and injects them as residuals into the DiT, sharing self‑attention key/value weights but using separate query/output weights.

Multi‑modal semantic conditioning : CLIP image embeddings of garment texture and textual attribute descriptions are injected via cross‑attention to improve consistency in occluded or sparsely reconstructed regions.

Training follows a progressive two‑stage strategy: first pre‑train on standard VVT data for stable garment transfer, then fine‑tune on synthetic multi‑view CaM‑VVT data to learn cross‑view consistency and realistic motion under arbitrary camera motions.

Why Simple Stacking Fails

A naïve pipeline that sequentially applies a video try‑on model and a separate camera‑control model suffers from error accumulation: unstable garment texture from the first stage is amplified when the second stage changes the viewpoint, leading to garment distortion, broken body structure, and background warping.

TryOnCrafter’s insight is that a stable intermediate representation—the renderable 4D proxy—is essential to bridge garment transfer and view control.

Experimental Results

Comparisons were made on both non‑camera‑controllable VVT and camera‑controllable VVT tasks. Visual results (see figures) demonstrate that TryOnCrafter produces structurally stable, texture‑consistent videos that follow arbitrary camera trajectories, outperforming prior methods.

Potential Applications

Subject re‑positioning: move the clothed person to any location in the scene while preserving occlusion and perspective.

Bullet‑time effect: freeze a pose and rotate the camera around the subject.

360° surround view: generate continuous views around the subject, enabling inspection of front, side, and back.

Robustness Tests

Two robustness axes were evaluated: foreground human errors and background geometry degradation. Results show that even with imperfect proxy reconstructions, the DiT can leverage source video references, garment semantics, and generative priors to plausibly fill missing or blurry regions.

Performance Tests

Full‑pipeline profiling reveals that building and rendering the 4D proxy takes 38.9 s (peak 59.1 GB GPU memory), while the DiT generation stage consumes ~1360 s (peak 68.6 GB). Re‑rendering the proxy for a new camera path costs only ~14.3 s (20.3 GB), indicating low marginal cost for multi‑view scenarios.

Future Outlook

Future work will focus on real‑time proxy construction and rendering, higher‑resolution view‑consistent generation, and longer sequence stability, aiming to deliver more intelligent, free‑form, immersive virtual try‑on experiences.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

SMPL-X3D Gaussian SplattingECCV 20264D try-on proxycamera controllablevideo diffusion transformervideo virtual try-on
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.