Paper Overview
Field: Computer Vision Authors: Hao Sun, Hao Yan, Mengting Chen Release Date: 2026-06-25 arXiv: 2606.19229
Abstract (Translated)
While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynamic subjects, existing paradigms remain fundamentally constrained by a passive dependency on source camera trajectories, failing to accommodate the requisite interactive freedom for omnidirectional viewpoint exploration. To address this limitation, the authors define a pioneering research frontier: Camera-controllable Video Virtual Try-on (CaM-VVT). Unlike conventional VVT, CaM-VVT not only necessitates viewpoint-agnostic texture hallucination but also strict structural synchronization between non-rigid human dynamics and background contexts under arbitrary, unconstrained camera movements.
To tackle these challenges, the authors present TryOnCrafter, the first unified DiT-based framework specifically built for the CaM-VVT task. Departing from implicit pixel-space manipulation, they introduce a renderable 4D try-on proxy that explicitly decouples the human subject from the environment. This is achieved by distilling high-fidelity 2D try-on priors into a 3DGS-based dressed avatar, which is then animated through SMPL-X sequences and metric-aligned to a reconstructed background point cloud. The proxy establishes a robust structural foundation with superior texture density and motion completeness.
The proxy-anchored video DiT leverages this robust structural foundation as the primary geometric anchor, ensuring that synthesized photorealistic videos strictly adhere to prescribed trajectories and physically plausible deformations. Owing to the inherent editability of the 4D proxy, TryOnCrafter facilitates diverse downstream applications, including human repositioning, "bullet-time" effects, and 360-degree orbit viewing.
--- *Automatically collected on 2026-06-26*