Paper Overview
Field: CV Authors: Ruozhen He, Meng Wei, Ziyan Yang, Vicente Ordonez Published: 2026-05-14 arXiv: 2605.15199
Original Abstract (translated)
Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and locations across shots remains a challenge over long sequences. Existing evaluations typically use independently generated prompt sets with limited entity coverage and simple consistency metrics, making standardized comparison difficult.
The authors introduce EntityBench, a benchmark of 140 episodes (2,491 shots) derived from real narrative media, with explicit per-shot entity schedules tracking characters, objects, and locations simultaneously across easy / medium / hard tiers of up to 50 shots, 13 cross-shot characters, 8 cross-shot locations, 22 cross-shot objects, and recurrence gaps spanning up to 48 shots. It is paired with a three-pillar evaluation suite (details truncated in the source post).
---
*Note: The Chinese summary included in the original forum post describes an unrelated paper on mixed-integer goal programming for dietary optimization and appears to be a mis-paired abstract. This page reflects the EntityBench paper indicated by the title and arXiv link.*
*Auto-collected on 2026-05-15*