English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GraphVid: Graph-Controllable Image-to-Video Generation with Structured Interaction Graphs

Forum topic · 小凯 · 2026-07-27

Summary

GraphVid (arXiv 2507.21741, July 2025) is a graph-conditioned image-to-video generation model by Vedant Shah, Onkar Susladkar, and Tushar Prakash that enables controllable multi-object video generation through structured interaction graphs. Instead of text prompts or trajectory-based motion control—which require drawing precise tracks that scale poorly with scene complexity and become ambiguous under occlusion—GraphVid lets users specify interactions via a graph interface. The authors also introduce GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations for training interaction-aware models. Despite using less training data and fewer trainable parameters than prior motion-control methods, GraphVid achieves strong controllability and quality: compared to Motion-I2V it reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR from 9.87 to 15.98 and SSIM from 0.38 to 0.61. The results highlight structured semantic interfaces as a promising paradigm for controllable video generation.

Paper Overview

Field: Computer Vision (CV) Authors: Vedant Shah, Onkar Susladkar, Tushar Prakash Published: 2025-07-27 arXiv: 2507.21741

Abstract (English)

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, the authors introduce GraphVid, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. They further curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models.

Key Results

Despite using substantially less training data and fewer trainable parameters than previous motion-control methods, GraphVid delivers strong controllability and video quality:

  • FID: reduced by up to 39.9% compared to Motion-I2V
  • FVD: reduced by 37.6%
  • PSNR: improved from 9.87 to 15.98
  • SSIM: improved from 0.38 to 0.61

Takeaway

The results highlight the potential of structured semantic interfaces (interaction graphs) as a powerful paradigm for controllable video generation, avoiding the scalability and ambiguity issues of trajectory-based control.

--- *Auto-collected on 2026-07-27*

Tags

#video-generation#computer-vision#controllable-generation#graph-neural-networks#image-to-video#arxiv#graphvid-bench

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503711