Paper Overview
Field: Computer Vision (CV) Authors: Qiaowei Miao, Kehan Li, Yawei Luo arXiv: 2507.00484
Abstract (Original)
Generative diffusion models excel at synthesizing high-quality images, videos, and 3D content under multimodal control. However, arbitrary user-defined modality-to-4D (X-to-4D) generation remains challenging due to the high cost of constructing diverse datasets and the limited scalability of existing methods. This paper presents Align4D, a flexible framework that translates any-modal input into coherent video-3D pairs, using video to guide 4D motion and 3D data to shape 4D geometry.
Align4D introduces three key techniques:
1. Object Distance Alignment: searches Video-Aligned and Multiview-Aligned Object Distances (VAOD/MAOD), respectively, to reconcile 4D renderings with video and the priors of multiview diffusion models. 2. Motion-Geometry Joint Alignment: constrains known and unknown views via synchronized video and 3D input constraints, ensuring consistent 4D generation. 3. Asynchronous Optimization: decouples Gaussian attribute and deformation network training to enhance motion and geometry fidelity.
Additionally, the authors propose the X4D dataset, integrating prompts, images, video, and 3D data for benchmarking.
Results
Experiments on X4D and Consistent4D demonstrate that Align4D achieves state-of-the-art quality and consistency in X-to-4D generation.
--- *Auto-collected on 2026-07-04. Source: zhichai.net forum post.*