English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Autoregression

Forum topic · 小凯 · 2026-07-08

Summary

MV-Forcing is a research paper by Gal Fiebelman, Hadar Averbuch-Elor, and Sagie Benaim (arXiv:2607.05376) addressing a longstanding gap in video diffusion models: generating long, multi-view consistent videos of dynamic scenes. While recent models can produce temporally autoregressive long single-view videos or short multi-view clips via bidirectional attention, combining both has remained unsolved. MV-Forcing unifies temporal and view autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. The key insight is that autoregressive 3D reconstruction models naturally interface between autoregressively generated views: given a completed source view, the method reconstructs its 3D structure and renders a geometric prior for the next target view, which the diffusion model then refines into high-quality video. The approach distills the model via distribution matching distillation and spatio-temporal self-forcing, closing the train-inference exposure bias of both temporal and view-wise sequential autoregression. Experiments show MV-Forcing generates geometrically consistent multi-view dynamic scene videos of arbitrary length and number of views using a single-step student model.

Paper Overview

Field: Computer Vision (CV) Authors: Gal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim Published: 2026-07-06 arXiv: 2607.05376

Abstract (English Translation)

Recent advances in video diffusion models have enabled generating long single-view videos through temporal autoregression, or short multi-view synthesis via bidirectional attention. However, generating long-range, multi-view consistent videos of dynamic scenes remains unsolved.

This paper proposes MV-Forcing, which combines temporal and view autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. The key insight is that autoregressive 3D reconstruction models naturally interface between autoregressively generated views: given a completed source view, the method reconstructs its 3D structure and renders a geometric prior for the next target view, which the diffusion model then refines into high-quality video.

The model is further distilled through distribution matching distillation and spatio-temporal self-forcing, closing the train-inference exposure bias of both temporal and view-wise sequential autoregression.

Experiments demonstrate that MV-Forcing generates geometrically consistent multi-view dynamic scene videos of arbitrary length and number of views using a single-step student model.

Links

  • arXiv paper: https://arxiv.org/abs/2607.05376
---

*Auto-collected on 2026-07-06*

Tags

#mv-forcing#video-generation#multi-view#diffusion-models#4d-reconstruction#autoregressive-models#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346206