Paper Overview
Field: Computer Vision (CV) Authors: Lee Hsin-Ying, Hanwen Jiang, Yiqun Mei Published: 2025-05-23 arXiv: 2505.17386
Abstract
Current motion-controlled image-to-video generation models strictly follow user-provided trajectories, but these trajectories are often sparse, imprecise, and causally incomplete. This dependence frequently produces unnatural or illogical results, notably omitting secondary causal consequences of motion.
To address this, the authors introduce MotiMotion, a new framework that reframes motion control as a "reason-then-generate" problem.
Key Contributions
- Training-free visual reasoning: A vision-language reasoner refines the primary trajectory's image-space coordinates and "imagines" plausible secondary motions, encouraging causally grounded, commonsense-consistent interactions.
- Confidence-aware control: A scheme modulates guidance strength so the model closely follows the plan when confidence is high, and corrects artifacts using its internal generative priors when confidence is low.
- MotiBench: A new image-to-video benchmark composed of interaction-centric scenes where novel events are triggered by motion, enabling systematic evaluation.
Results
Both VLM-based evaluations and human studies on MotiBench show that MotiMotion generates videos with more plausible object behavior and interactions, outperforming existing methods.
---
*Auto-collected on 2026-05-23*