Paper Overview
Research Area: CV Authors: Xiang Fan, Yuheng Wang, Bohan Fang, Zhongzheng Ren, Ranjay Krishna Published: 2026-05-14 arXiv: 2605.15196
Forum Summary
The original forum post includes an abstract (in Chinese) discussing multi-agent orchestration and orchestrator invisibility, which appears inconsistent with the linked paper's title and abstract. The original English abstract below, matching the RefDecoder paper, is the authoritative source.
Original Abstract
Video generation powers a vast array of downstream applications. However, while the de facto standard, i.e., latent diffusion models, typically employ heavily conditioned denoising networks, their decoders often remain unconditional. We observe that this architectural asymmetry leads to significant loss of detail and inconsistency relative to the input image. To address this, we argue that the decoder requires equal conditioning to preserve structural integrity. We introduce RefDecoder, a reference-conditioned video VAE decoder by injecting high-fidelity reference image signal directly into the decoding process via reference attention. Specifically, a lightweight image encoder maps the reference frame into the detail-rich high-dimensional tokens, which are co-processed with the denoised video latents.
Notes
- Full abstract text is truncated in the source post; refer to the arXiv page for the complete version.
- Auto-collected on 2026-05-15.