English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

TryOnCrafter: Camera-Controllable Video Virtual Try-on via a Renderable 4D Try-on Proxy

Forum topic · 小凯 · 2026-06-26

Summary

TryOnCrafter (arXiv 2606.19229) introduces Camera-controllable Video Virtual Try-on (CaM-VVT), a new research frontier addressing the limitation of existing Video Virtual Try-on (VVT) systems, which passively depend on source camera trajectories and cannot support omnidirectional viewpoint exploration. The proposed unified DiT-based framework uses a renderable 4D try-on proxy that explicitly decouples the human subject from the environment: high-fidelity 2D try-on priors are distilled into a 3DGS-based dressed avatar, animated via SMPL-X sequences and metric-aligned with a reconstructed background point cloud. This proxy serves as a robust geometric anchor, ensuring synthesized videos strictly follow specified camera trajectories while producing physically plausible deformations and view-consistent textures. Thanks to the proxy's editability, TryOnCrafter enables downstream applications including human repositioning, bullet-time effects, and 360-degree orbit viewing.

Paper Overview

Field: Computer Vision Authors: Hao Sun, Hao Yan, Mengting Chen Release Date: 2026-06-25 arXiv: 2606.19229

Abstract (Translated)

While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynamic subjects, existing paradigms remain fundamentally constrained by a passive dependency on source camera trajectories, failing to accommodate the requisite interactive freedom for omnidirectional viewpoint exploration. To address this limitation, the authors define a pioneering research frontier: Camera-controllable Video Virtual Try-on (CaM-VVT). Unlike conventional VVT, CaM-VVT not only necessitates viewpoint-agnostic texture hallucination but also strict structural synchronization between non-rigid human dynamics and background contexts under arbitrary, unconstrained camera movements.

To tackle these challenges, the authors present TryOnCrafter, the first unified DiT-based framework specifically built for the CaM-VVT task. Departing from implicit pixel-space manipulation, they introduce a renderable 4D try-on proxy that explicitly decouples the human subject from the environment. This is achieved by distilling high-fidelity 2D try-on priors into a 3DGS-based dressed avatar, which is then animated through SMPL-X sequences and metric-aligned to a reconstructed background point cloud. The proxy establishes a robust structural foundation with superior texture density and motion completeness.

The proxy-anchored video DiT leverages this robust structural foundation as the primary geometric anchor, ensuring that synthesized photorealistic videos strictly adhere to prescribed trajectories and physically plausible deformations. Owing to the inherent editability of the 4D proxy, TryOnCrafter facilitates diverse downstream applications, including human repositioning, "bullet-time" effects, and 360-degree orbit viewing.

--- *Automatically collected on 2026-06-26*

Tags

#computer-vision#video-virtual-try-on#generative-models#3dgs#smpl-x#camera-control#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208132