English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

Forum topic · 小凯 · 2026-05-23

Summary

MotiMotion is a new framework for image-to-video generation that reframes motion control as a reasoning-then-generation problem. Existing motion-controlled models strictly follow user-provided trajectories, which are often sparse, imprecise, and causally incomplete, leading to unnatural or illogical results—especially missing secondary causal consequences of motion. MotiMotion addresses this by using training-free vision-language reasoning to refine primary trajectory image-space coordinates and to 'imagine' plausible secondary motions, encouraging causally grounded and commonsense-consistent interactions. A confidence-aware control scheme modulates guidance strength: the model follows the plan closely when confidence is high and corrects artifacts with its internal generative priors when confidence is low. The authors also introduce MotiBench, a new image-to-video benchmark of interaction-centric scenes where novel events are triggered by motion. VLM-based evaluations and human studies on MotiBench show that MotiMotion produces videos with more plausible object behavior and interactions, outperforming existing methods. Paper: arXiv:2505.17386.

Paper Overview

Field: Computer Vision (CV) Authors: Lee Hsin-Ying, Hanwen Jiang, Yiqun Mei Published: 2025-05-23 arXiv: 2505.17386

Abstract

Current motion-controlled image-to-video generation models strictly follow user-provided trajectories, but these trajectories are often sparse, imprecise, and causally incomplete. This dependence frequently produces unnatural or illogical results, notably omitting secondary causal consequences of motion.

To address this, the authors introduce MotiMotion, a new framework that reframes motion control as a "reason-then-generate" problem.

Key Contributions

  • Training-free visual reasoning: A vision-language reasoner refines the primary trajectory's image-space coordinates and "imagines" plausible secondary motions, encouraging causally grounded, commonsense-consistent interactions.
  • Confidence-aware control: A scheme modulates guidance strength so the model closely follows the plan when confidence is high, and corrects artifacts using its internal generative priors when confidence is low.
  • MotiBench: A new image-to-video benchmark composed of interaction-centric scenes where novel events are triggered by motion, enabling systematic evaluation.

Results

Both VLM-based evaluations and human studies on MotiBench show that MotiMotion generates videos with more plausible object behavior and interactions, outperforming existing methods.

---

*Auto-collected on 2026-05-23*

Tags

#video-generation#computer-vision#motion-control#visual-reasoning#vlm#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620657