[论文] Never Look Back: Understanding Persistence in 3D Object Memory from Eg...

研究领域: CV 作者: Shravan Chaudhari, William Paul, Suchi Saria, Rama Chellappa, Homanga Bharadhwaj 发布时间: 2026-10-07 arXiv: 2610.10538

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Shravan Chaudhari, William Paul, Suchi Saria, Rama Chellappa, Homanga Bharadhwaj 发布时间: 2026-10-07 arXiv: 2610.10538

中文摘要

当我们在世界中移动并完成日常任务时,会接触到一些物体,它们可能只在之后才变得重要。我们能够回忆起把东西放在了哪里、容器里装了什么,即使事先并不知道自己会需要这些信息。本研究探讨了具身助手如何从第一人称视频中构建类似的记忆。我们提出Ledger——一个持久的3D物体记忆系统,整合物体位置、历史轨迹和上下文描述。它将整个记录中的观察关联起来,在物体离开视野后仍保留其信息——包括那些使用者从未触碰过的物体。系统按静置位置对每个物体的观察进行聚类,仅在获得重复证据后才记录移动,从而降低定位噪声的影响。简短描述保留了物体内容或支撑面等细节。这些记录被保存下来,之后无需访问原始图像或视频即可回答空间问题。我们的记忆系统将HD-EPIC准确率从29.7%提升至42.6%,UCS-Bench准确率从33.8%提升至38.5%,在Ego4D物体定位任务上返回预测的中位误差为0.99米。

原文摘要

As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise....


*自动采集于 2026-10-09*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens