English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Can 4D Foundation Models Remember? PersistBench Evaluates Visual Memory

Forum topic · 小凯 · 2026-09-19

Summary

PersistBench is a new dataset and metric suite for evaluating how well 4D foundation models—including camera-controllable video models and 4D reconstruction models—remember what they perceive. Existing benchmarks rely on pixel-level metrics and lack ground truth for objects once they leave the field of view, making object-centric evaluation of visual memory impossible. PersistBench addresses this by using 360° videos as omniscient ground truth and introduces three evaluation aspects: object permanence, motion continuity, and appearance preservation. Experiments across multiple model categories reveal that current models maintain only short-term consistency and performance drops significantly once objects exit the field of view. The authors conclude that 'seeing is not remembering,' highlighting a gap between current model capabilities and robust visual memory. The paper is authored by Guangzhao He, Hadar Averbuch-Elor, and Wei-Chiu Ma (arXiv:2609.20819), with dataset and code available at https://guangzhaohe.com/persistbench.

Paper Overview

  • Field: Computer Vision
  • Authors: Guangzhao He, Hadar Averbuch-Elor, Wei-Chiu Ma
  • arXiv: 2609.20819
  • Motivation

    Perceiving and remembering the visual world is fundamental to navigating and interacting with our environment. Current 4D foundation models—such as camera-controllable video models or 4D reconstruction models—can perceive and reconstruct dynamic environments, but how well they remember what they have perceived remains an open question.

    Existing benchmarks largely rely on pixel-level metrics and lack ground truth for objects once they leave the field of view, making them unable to evaluate visual memory in an object-centric manner against references.

    PersistBench

    To fill this gap, the authors introduce PersistBench, a dataset and metric suite that:

  • Leverages 360° videos as omniscient ground truth, enabling reference-based evaluation even for objects outside the model's field of view.
  • Proposes three evaluation aspects:
  • 1. Object permanence — whether objects are remembered to still exist after leaving view 2. Motion continuity — whether object motion remains consistent over time 3. Appearance preservation — whether object appearance is faithfully maintained

    Findings

    Evaluating a variety of models across multiple categories, the study finds that current models maintain only short-term consistency; performance degrades significantly once objects leave the field of view. The results highlight the gap between current model capabilities and robust visual memory—"seeing is not remembering"—providing guidance for the future development of 4D foundation models.

    Resources

  • Paper: https://arxiv.org/abs/2609.20819
  • Dataset and code: https://guangzhaohe.com/persistbench

Tags

#computer-vision#4d-reconstruction#video-models#benchmark#visual-memory#persistbench#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634974