Summary
GraphVid is a graph-conditioned image-to-video generation model introduced to overcome the difficulty of specifying precise multi-object interactions via text prompts or trajectory-based motion controls, which scale poorly with scene complexity and become ambiguous under occlusion. Instead of pixel-level trajectories, GraphVid enables interactive control through structured interaction graphs that describe relations between subjects. The authors also release GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations for training interaction-aware video generation models. Despite using less training data and fewer trainable parameters than prior motion-control approaches, GraphVid achieves strong controllability and video quality: compared to Motion-I2V, it reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR from 9.87 to 15.98 and SSIM from 0.38 to 0.61. The work highlights structured semantic interfaces as a powerful paradigm for controllable video generation. Paper: arXiv:2507.19315.
Overview
- Field: Computer Vision (CV)
- Authors: Vedant Shah, Onkar Susladkar, Tushar Prakash
- Published: 2026-07-24
- arXiv: 2507.19315
Key points
- Controllable video generation remains challenging: text prompts or motion-control inputs that primarily constrain pixel movement make it hard to specify precise multi-object interactions.
- Trajectory-based control requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap.
- GraphVid is a graph-conditioned image-to-video generation model enabling interactive control through structured interaction graphs, offering flexible yet precise multi-subject control.
- The authors curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to support training of interaction-aware video generation models.
- Despite using far less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality.
Results
Compared to Motion-I2V:
- FID reduced by up to 39.9%
- FVD reduced by 37.6%
- PSNR improved: 9.87 → 15.98
- SSIM improved: 0.38 → 0.61
The results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.
---
*Auto-collected on 2026-07-25.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178447085