[论文] Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity...
研究领域: NLP 作者: Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua 发布时间: 2026-09-29 arXiv: 2609.38155
论文概要
研究领域: NLP 作者: Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua 发布时间: 2026-09-29 arXiv: 2609.38155
中文摘要
回答关于长视频的问题往往需要跨数小时甚至数天连接同一物体的事件。按时间顺序的描述和文本推导的实体可能无法解决物理身份问题:不同物体可能共享同一描述,而同一物体的观测在不同事件中仍然断开。检索相关事件因此不一定能恢复问题所关注特定实体的'传记'。为此,我们引入GEB(Grounded Entity Biographies)——一个长视频记忆框架,将跨片段的视觉锚定观测分组为同一物理实例的可检索传记,同时保留每个时刻的上下文。在问答过程中,传记与情节证据一起被检索,使模型能够利用记忆构建期间建立的身份链接来追踪实体在各事件中的轨迹。在四个基准上的评估——包括跨天和跨周的录制——表明该方法在选择题和开放式问答上均优于先前的记忆框架。在EgoLifeQA上,GEB达到72.0%的准确率,比已发表的最佳结果高4.4个百分点。消融实验表明,锚定的身份关联和传记阅读都对性能提升有贡献,这是仅靠额外描述无法完全恢复的。
原文摘要
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allow...
*自动采集于 2026-10-01*
#论文 #arXiv #NLP #小凯