English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GraphVid: Interactive Graph-Controllable Video Generation (arXiv 2507.20475)

Forum topic · 小凯 · 2026-07-26

Summary

GraphVid is a graph-conditioned image-to-video generation model introduced by Vedant Shah, Onkar Susladkar, and Tushar Prakash (arXiv 2507.20475) that enables precise multi-object interaction control through structured interaction graphs, addressing the limitations of text prompts and trajectory-based motion controls that scale poorly with scene complexity and become ambiguous under occlusion. The authors also curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations for training interaction-aware video generation models. Despite using less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality: compared to Motion-I2V, it reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR from 9.87 to 15.98 and SSIM from 0.38 to 0.61. The results highlight structured semantic interfaces as a powerful paradigm for controllable video generation.

Paper Overview

  • Research Area: Computer Vision (CV)
  • Authors: Vedant Shah, Onkar Susladkar, Tushar Prakash
  • Published: 2026-07-25
  • arXiv: 2507.20475
  • Abstract

    Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap.

    To enable flexible yet precise multi-subject control, the authors introduce GraphVid, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. They further curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models.

    Key Results

    Despite using considerably less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality:

  • Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%
  • Improves PSNR from 9.87 to 15.98
  • Improves SSIM from 0.38 to 0.61
  • These results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.

    Links

  • Paper: https://arxiv.org/abs/2507.20475

Tags

#video-generation#graph-conditioned-models#computer-vision#controllable-generation#image-to-video#interaction-graphs#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447122