English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

RefDecoder: Enhancing Visual Generation with Reference-Conditioned Video Decoding

Forum topic · 小凯 · 2026-05-15

Summary

RefDecoder (arXiv:2605.15196, Xiang Fan, Yuheng Wang, Bohan Fang, Zhongzheng Ren, Ranjay Krishna) addresses a key architectural asymmetry in latent diffusion models for video generation: while denoising networks are heavily conditioned, their VAE decoders typically remain unconditional, causing significant detail loss and inconsistency relative to the input image. The authors argue that the decoder requires equal conditioning to preserve structural integrity. RefDecoder is a reference-conditioned video VAE decoder that injects high-fidelity reference image signal directly into the decoding process via reference attention. A lightweight image encoder maps the reference frame into detail-rich high-dimensional tokens, which are co-processed with the denoised video latents during decoding. Posted on the zhichai.net forum on 2026-05-15, the paper targets the computer vision community and proposes a drop-in decoder improvement for video generation pipelines.

Paper Overview

Research Area: CV Authors: Xiang Fan, Yuheng Wang, Bohan Fang, Zhongzheng Ren, Ranjay Krishna Published: 2026-05-14 arXiv: 2605.15196

Forum Summary

The original forum post includes an abstract (in Chinese) discussing multi-agent orchestration and orchestrator invisibility, which appears inconsistent with the linked paper's title and abstract. The original English abstract below, matching the RefDecoder paper, is the authoritative source.

Original Abstract

Video generation powers a vast array of downstream applications. However, while the de facto standard, i.e., latent diffusion models, typically employ heavily conditioned denoising networks, their decoders often remain unconditional. We observe that this architectural asymmetry leads to significant loss of detail and inconsistency relative to the input image. To address this, we argue that the decoder requires equal conditioning to preserve structural integrity. We introduce RefDecoder, a reference-conditioned video VAE decoder by injecting high-fidelity reference image signal directly into the decoding process via reference attention. Specifically, a lightweight image encoder maps the reference frame into the detail-rich high-dimensional tokens, which are co-processed with the denoised video latents.

Notes

  • Full abstract text is truncated in the source post; refer to the arXiv page for the complete version.
  • Auto-collected on 2026-05-15.

Tags

#video-generation#diffusion-models#vae#decoder#reference-attention#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620058