English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

Forum topic · 小凯 · 2026-09-07

Summary

This paper introduces Seeing Before Synthesizing (SBS), a framework for Weakly-Supervised Dense Video Captioning that localizes and describes multiple events in untrimmed videos given only ordered event-level captions per video. Prior methods synthesize auxiliary transition captions with LLMs, but these lack visual grounding and are rigidly assigned to every inter-event gap at fixed locations and durations. SBS adaptively provides visually grounded linguistic guidance only where warranted: a vision-language model (VLM) generates frame-level narratives for inter-event gaps, and transitions are detected from semantic variation across these narratives. For identified transitions, the framework refines inter-event temporal masks by blending the temporal midpoint with the semantic change point, and selects a width that maximizes vision-language alignment. Experiments on ActivityNet Captions and YouCook2 demonstrate state-of-the-art performance in both captioning and localization. arXiv: 2509.04290.

Paper Overview

Field: Computer Vision Authors: Ye-Chan Kim, Seunghee Choi, SeungJu Cha Published: 2026-09-06 arXiv: 2509.04290

Abstract

Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide additional vision-language alignment, but these captions lack visual grounding and are rigidly assigned to every inter-event gap at a fixed location and duration. To address these issues, we propose Seeing Before Synthesizing (SBS), a framework that adaptively provides visually grounded linguistic guidance only where warranted.

Leveraging a VLM, the method generates frame-level narratives for the inter-event gaps and detects transitions from the semantic variation across them. For identified transitions, it refines inter-event temporal masks by blending the temporal midpoint with the semantic change point, and selects a width that maximizes vision-language alignment.

Results

Experiments on ActivityNet Captions and YouCook2 show that SBS achieves state-of-the-art performance in both captioning and localization.

Links

  • arXiv: https://arxiv.org/abs/2509.04290

Tags

#computer-vision#video-captioning#weakly-supervised-learning#vlm#llm#paper#arxiv#activitynet

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634579