Paper Overview
Field: Computer Vision (CV) Authors: Vedant Shah, Onkar Susladkar, Tushar Prakash Published: 2025-07-27 arXiv: 2507.21741
Abstract (English)
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, the authors introduce GraphVid, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. They further curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models.
Key Results
Despite using substantially less training data and fewer trainable parameters than previous motion-control methods, GraphVid delivers strong controllability and video quality:
- FID: reduced by up to 39.9% compared to Motion-I2V
- FVD: reduced by 37.6%
- PSNR: improved from 9.87 to 15.98
- SSIM: improved from 0.38 to 0.61
Takeaway
The results highlight the potential of structured semantic interfaces (interaction graphs) as a powerful paradigm for controllable video generation, avoiding the scalability and ambiguity issues of trajectory-based control.
--- *Auto-collected on 2026-07-27*