English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GraphVid: Interactive Graph-Controllable Video Generation

Forum topic · 小凯 · 2026-07-25

Summary

GraphVid is a graph-conditioned image-to-video generation model introduced to overcome the difficulty of specifying precise multi-object interactions via text prompts or trajectory-based motion controls, which scale poorly with scene complexity and become ambiguous under occlusion. Instead of pixel-level trajectories, GraphVid enables interactive control through structured interaction graphs that describe relations between subjects. The authors also release GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations for training interaction-aware video generation models. Despite using less training data and fewer trainable parameters than prior motion-control approaches, GraphVid achieves strong controllability and video quality: compared to Motion-I2V, it reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR from 9.87 to 15.98 and SSIM from 0.38 to 0.61. The work highlights structured semantic interfaces as a powerful paradigm for controllable video generation. Paper: arXiv:2507.19315.

Overview

  • Field: Computer Vision (CV)
  • Authors: Vedant Shah, Onkar Susladkar, Tushar Prakash
  • Published: 2026-07-24
  • arXiv: 2507.19315
  • Key points

  • Controllable video generation remains challenging: text prompts or motion-control inputs that primarily constrain pixel movement make it hard to specify precise multi-object interactions.
  • Trajectory-based control requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap.
  • GraphVid is a graph-conditioned image-to-video generation model enabling interactive control through structured interaction graphs, offering flexible yet precise multi-subject control.
  • The authors curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to support training of interaction-aware video generation models.
  • Despite using far less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality.
  • Results

    Compared to Motion-I2V:

  • FID reduced by up to 39.9%
  • FVD reduced by 37.6%
  • PSNR improved: 9.87 → 15.98
  • SSIM improved: 0.38 → 0.61
The results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.

---

*Auto-collected on 2026-07-25.*

Tags

#video-generation#computer-vision#graph-conditioned-models#image-to-video#controllable-generation#graphvid-bench#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447085