Loading...
正在加载...
请稍候

[论文] PercepCap: Video Captioner with Structured Spatio-Temporal Perception

小凯 (C3P0) 2026年07月24日 00:46

论文概要

研究领域: CV
作者: Yifan Xu, Zihao Wang, Zhixiao Wang
发布时间: 2026-07-24
arXiv: 2507.18394

中文摘要

视频字幕生成需要对视频进行细粒度的时空理解,包括物体所在位置的空间感知和事件何时发生的时间感知。现有的MLLM通常直接从视频输入生成字幕,而不暴露描述背后的感知证据。因此,时空感知中的错误仅在最终字幕中观察到,使得直接识别底层感知错误变得困难。为解决这些问题,我们提出了PercepCap,一种感知感知视频字幕生成框架,在生成最终字幕之前使感知证据显式化。具体而言,PercepCap遵循感知-描述生成链,模型首先产生包含物体轨迹和时间事件的时空感知痕迹,然后基于感知证据生成最终字幕。为支持这一点,我们设计了两阶段训练策略。感知-然后-描述监督微调使模型从仅字幕生成适应到提出的感知-描述链,而感知基础强化学习通过感知链和最终字幕的联合奖励优化感知痕迹和字幕质量。为支持我们的两阶段训练,我们引入了字幕锚定感知数据构建。该管道通过首先生成仅字幕描述,提取其中提到的物体和事件,并用框和时间戳将它们接地回视频来构建SFT和RL训练数据。这产生了字幕对齐的感知数据,提供了可靠的训练真值,确保显式感知痕迹和最终字幕引用相同的物体和事件。在直接字幕和字幕到QA评估中,PercepCap始终优于Qwen3-VL基线,并展示了领先的字幕质量。

原文摘要

Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occur. Existing MLLMs usually generate captions directly from video inputs without exposing the perceptual evidence behind descriptions. As a result, mistakes in spatiotemporal perception are only observed in the final caption, making it difficult to identify the underlying perceptual errors directly. To address these issues, we present PercepCap, a perception-aware video captioning framework that makes perceptual evidence explicit before producing the final caption. Specifically, PercepCap follows a perceive-describe generation chain, where the model first produces a spatiotemporal perception trace comprising objec...


自动采集于 2026-07-24

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录