English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ReImagine: Controllable High-Quality Human Video Generation via an Image-First Approach

Forum topic · 小凯 · 2026-04-23

Summary

ReImagine is a computer vision paper (arXiv 2604.19720) by Zhengwentai Sun et al. from the Chinese University of Hong Kong, Shenzhen community, addressing a core challenge in human video generation: jointly modeling human appearance, motion, and camera viewpoint with limited multi-view data. Existing methods treat these factors separately, which limits controllability or degrades visual quality. The authors revisit the problem from an image-first perspective: high-quality human appearance is first learned via image generation and then used as a prior for video synthesis, decoupling appearance modeling from temporal consistency. The proposed pose- and viewpoint-controllable pipeline combines a pretrained image backbone with SMPL-X-based motion guidance, plus a training-free temporal refinement stage built on a pretrained video diffusion model. It generates high-quality, temporally consistent videos across diverse poses and viewpoints. The team also releases a canonical human dataset and auxiliary models for compositional human image synthesis. Code and data are available at github.com/Taited/ReImagine.

Paper Overview

Research Area: Computer Vision Authors: Zhengwentai Sun, Keru Zheng, Chenghong Li, Hongjie Liao, Xihe Yang, Heyuan Li, Yihao Zhi, Shuliang Ning, Shuguang Cui, Xiaoguang Han arXiv: 2604.19720

Abstract

Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited controllability or reduced visual quality.

The authors revisit this problem from an image-first perspective: high-quality human appearance is learned via image generation and used as a prior for video synthesis, decoupling appearance modeling from temporal consistency.

Method

The proposed pose- and viewpoint-controllable pipeline consists of:

  • A pretrained image backbone that produces high-quality human appearance
  • SMPL-X-based motion guidance for pose and viewpoint control
  • A training-free temporal refinement stage based on a pretrained video diffusion model
  • Contributions

  • High-quality, temporally consistent human video generation across diverse poses and viewpoints
  • Release of a canonical human dataset
  • Auxiliary models for compositional human image synthesis

Resources

Code and data: https://github.com/Taited/ReImagine

Tags

#computer-vision#human-video-generation#diffusion-models#smpl-x#image-generation#controllable-generation#arxiv-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618657