English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction

Forum topic · 小凯 · 2026-06-14

Summary

Flex4DHuman is a multi-view video diffusion model that converts monocular or sparse multi-view videos of humans into synchronized, dense multi-view videos using only relative camera-pose conditioning. Built on Wan 2.1 1.3B and trained with a mixed-data strategy, the model achieves zero-shot generalization to animal categories beyond humans. When combined with 4D Gaussian Splatting, Flex4DHuman lifts single-view footage into dynamic 4D Gaussian representations suitable for AR/VR applications and simulation. The paper, authored by Jen-Hao Cheng, Yipeng Wang, Hao Zhang, Gengshan Yang, and Jenq-Neng Hwang, is available on arXiv as 2606.13655. This approach addresses a key bottleneck in 4D human capture: generating consistent novel views from limited camera coverage without requiring dense multi-camera rigs.

Overview

Field: Computer Vision Authors: Jen-Hao Cheng, Yipeng Wang, Hao Zhang, Gengshan Yang, Jenq-Neng Hwang Published: 2026-06-11 arXiv: 2606.13655

Abstract

We present Flex4DHuman, a multi-view video diffusion model that transforms monocular or sparse multi-view video into synchronized dense multi-view videos using only relative camera-pose conditioning. Built on Wan 2.1 1.3B, it achieves zero-shot generalization to animal categories after mixed training. Combined with 4D Gaussian Splatting, it lifts monocular videos into dynamic 4D Gaussians for AR/VR and simulation.

Key Highlights

  • Pose-only conditioning: Generates dense multi-view video from monocular or sparse inputs conditioned solely on relative camera poses.
  • Backbone: Built on Wan 2.1 1.3B video diffusion model.
  • Zero-shot generalization: Mixed training enables generalization to animal categories without dedicated fine-tuning.
  • 4D pipeline: Integration with 4D Gaussian Splatting lifts monocular video into dynamic 4D Gaussians.
  • Applications: Targets AR/VR content creation and simulation.

Tags

#computer-vision#diffusion-models#4d-reconstruction#gaussian-splatting#human-pose#video-generation#ar-vr

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981284