English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation

Forum topic · 小凯 · 2026-05-09

Summary

ActCam is a training-free (zero-shot) video generation method that enables fine-grained control over both actor motion and cinematography. It jointly transfers character motion from a driving video into a new scene while allowing per-frame control of intrinsic and extrinsic camera parameters. Built on any pretrained image-to-video diffusion model that accepts scene depth and character pose conditioning, ActCam generates geometrically consistent pose and depth conditions across frames. It uses a single sampling process with a two-phase conditioning schedule: early denoising steps condition on both pose and sparse depth to reinforce scene structure, after which depth is dropped so pose alone refines high-frequency details without over-constraining generation. Benchmarks covering diverse character motions and challenging viewpoint changes show improved camera adherence and action fidelity over pose-only baselines and other joint pose-and-camera methods, with human evaluators preferring ActCam especially under large viewpoint changes.

Paper: ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation Field: Computer Vision Authors: Omar El Khalifi, Thomas Rossi, Oscar Fossey arXiv: 2505.03486

Overview

For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. ActCam is a zero-shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per-frame control of intrinsic and extrinsic camera parameters.

Method

  • ActCam builds on any pretrained image-to-video diffusion model that accepts conditioning in terms of scene depth and character pose.
  • Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames.
  • Sampling runs in a single pass with a two-phase conditioning schedule:
  • Early denoising steps: condition on both pose and sparse depth to strengthen scene structure.
  • Later steps: drop depth and guide with pose only, refining high-frequency details and avoiding over-constraining the generative process.
  • Results

  • Evaluated on multiple benchmarks spanning diverse character actions and challenging viewpoint changes.
  • Compared against pose-only control and other joint pose-and-camera methods, ActCam improves both camera adherence and action fidelity.
  • Human evaluators prefer ActCam, especially in scenes with large viewpoint changes.
  • Conclusion: careful camera-consistent conditioning plus staged guidance enables strong joint camera and motion control without any training.
--- *Collected automatically on 2026-05-09*

Tags

#video-generation#diffusion-models#camera-control#3d-motion#computer-vision#zero-shot#pose-transfer

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619660