Summary
MV-Forcing is a research paper by Gal Fiebelman, Hadar Averbuch-Elor, and Sagie Benaim (arXiv:2607.05376) addressing a longstanding gap in video diffusion models: generating long, multi-view consistent videos of dynamic scenes. While recent models can produce temporally autoregressive long single-view videos or short multi-view clips via bidirectional attention, combining both has remained unsolved. MV-Forcing unifies temporal and view autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. The key insight is that autoregressive 3D reconstruction models naturally interface between autoregressively generated views: given a completed source view, the method reconstructs its 3D structure and renders a geometric prior for the next target view, which the diffusion model then refines into high-quality video. The approach distills the model via distribution matching distillation and spatio-temporal self-forcing, closing the train-inference exposure bias of both temporal and view-wise sequential autoregression. Experiments show MV-Forcing generates geometrically consistent multi-view dynamic scene videos of arbitrary length and number of views using a single-step student model.
Paper Overview
Field: Computer Vision (CV)
Authors: Gal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim
Published: 2026-07-06
arXiv: 2607.05376
Abstract (English Translation)
Recent advances in video diffusion models have enabled generating long single-view videos through temporal autoregression, or short multi-view synthesis via bidirectional attention. However, generating long-range, multi-view consistent videos of dynamic scenes remains unsolved.
This paper proposes MV-Forcing, which combines temporal and view autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. The key insight is that autoregressive 3D reconstruction models naturally interface between autoregressively generated views: given a completed source view, the method reconstructs its 3D structure and renders a geometric prior for the next target view, which the diffusion model then refines into high-quality video.
The model is further distilled through distribution matching distillation and spatio-temporal self-forcing, closing the train-inference exposure bias of both temporal and view-wise sequential autoregression.
Experiments demonstrate that MV-Forcing generates geometrically consistent multi-view dynamic scene videos of arbitrary length and number of views using a single-step student model.
Links
- arXiv paper: https://arxiv.org/abs/2607.05376
---
*Auto-collected on 2026-07-06*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178346206