[论文] Can 4D Foundation Models Remember?

研究领域: CV 作者: Guangzhao He, Hadar Averbuch-Elor, Wei-Chiu Ma 发布时间: 2026-09-17 arXiv: 2609.20819

论文概要

研究领域: CV 作者: Guangzhao He, Hadar Averbuch-Elor, Wei-Chiu Ma 发布时间: 2026-09-17 arXiv: 2609.20819

中文摘要

感知和记忆视觉世界是导航和与环境交互的基础。当前的 4D 基础模型(如相机可控视频模型或 4D 重建模型)能够感知和重建动态环境,但它们对所感知内容的记忆能力如何仍是一个悬而未决的问题。现有基准大多依赖像素级指标,且缺乏物体离开视野后的真值,因此无法以物体为中心、对照参考来评估视觉记忆。为填补这一空白,我们提出 PersistBench——一个数据集和指标套件,利用 360° 视频作为全知真值,提出三个评估维度:物体恒存性、运动连续性和外观保持性。跨多个类别评估各种模型后发现,当前模型只能维持短期一致性,一旦物体离开视野,性能就会显著下降。我们的发现凸显了当前模型能力与稳健视觉记忆之间的差距("看见不等于记住"),为 4D 基础模型的未来发展提供了指导。数据集和代码可在项目页面获取:https://guangzhaohe.com/persistbench。

原文摘要

Perceiving and remembering the visual world is fundamental to navigating and interacting with our environment. Current 4D foundation models, such as camera-controllable video models or 4D reconstruction models, can perceive and reconstruct dynamic environments, but how well they remember what they have perceived remains an open question. Existing benchmarks largely rely on pixel-level metrics and lack ground truth for objects once they leave the field of view, making them unable to evaluate visual memory in an object-centric manner against references. To fill this gap, we introduce PersistBench, a dataset and metric suite that leverages 360° videos as omniscient ground truth and proposes three evaluation aspects: object permanence, motion continuity, and appearance preservation. Evaluating...


*自动采集于 2026-09-19*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens