Summary
Warp-as-History (arXiv:2605.15182) is a computer vision paper by Yifan Wang and Tong He proposing a simple interface for generalizable camera-controlled video generation. Existing methods typically learn camera conditioning via camera encoders, control branches, or modified attention and positional encodings, requiring costly post-training on large-scale camera-annotated videos. Training-free alternatives avoid post-training but shift the cost to test-time optimization or extra denoising-time guidance. The proposed approach instead converts camera-induced warps into camera-warped pseudo-history with target-frame positional alignment and visible-token selection: given a target camera trajectory, the model constructs pseudo-history frames by warping past observations to the target viewpoint, conditioning generation without additional training or guidance overhead. Posted on zhichai.net's paper channel, the article links to the arXiv abstract and invites discussion of the method's generalization across trajectories and scenes.
Paper Overview
Field: Computer Vision (CV)
Authors: Yifan Wang, Tong He
Published: 2026-05-14
arXiv: 2605.15182
Abstract
Camera-controlled video generation has made substantial progress, enabling generated videos to follow prescribed viewpoint trajectories. However, existing methods usually learn camera-specific conditioning through camera encoders, control branches, or attention and positional-encoding modifications, which often require post-training on large-scale camera-annotated videos. Training-free alternatives avoid such post-training, but often shift the cost to test-time optimization or extra denoising-time guidance.
We propose Warp-as-History, a simple interface that turns camera-induced warps into camera-warped pseudo-history with target-frame positional alignment and visible-token selection. Given a target camera trajectory, we construct camera-warped pseudo-history from past observations and feed it to the video diffusion model as history context, allowing camera control without dedicated post-training or test-time guidance.
> Note: The abstract is truncated in the source post; see the arXiv page for the full text.
Discussion
Feel free to discuss the approach below — particularly how it compares to camera-encoder-based conditioning methods.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177620063