English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MemoryVLA++: Temporal Modeling via Memory and Imagination for Vision-Language-Action Models

Forum topic · 小凯 · 2026-06-10

Summary

MemoryVLA++ (arXiv:2506.04876) is a temporal modeling framework for vision-language-action (VLA) models in robotic manipulation, inspired by human cognitive mechanisms: working memory, hippocampal episodic memory, and internal models for imagining future states. A pretrained vision-language model encodes the current observation into perceptual and cognitive tokens that form working memory. These tokens query a perception-cognition memory bank that stores low-level details and high-level semantics of past interactions, updated via redundancy-aware consolidation. A world model imagines future states in a denoising latent space, and the imagined latents are integrated into temporally aware tokens under memory guidance. These tokens condition a diffusion action expert to predict temporally consistent action sequences. Experiments on 5 simulation benchmarks and 3 real-world robot task categories show improvements of +9% on general manipulation, +26% on memory-dependent tasks, and +28% on imagination-dependent tasks over baselines, addressing the weakness of most VLA models that rely only on the current observation.

Paper Overview

Field: Computer Vision / Robotics Authors: Hao Shi, Weiye Li, Bin Xie Published: 2025-06-06 arXiv: 2506.04876

Summary

Temporal modeling is essential for robotic manipulation: effective control requires both memory of past interactions and imagination of future states. However, most VLA (Vision-Language-Action) models rely primarily on the current observation and therefore struggle on long-horizon, temporally dependent tasks.

Cognitive science suggests humans rely on working memory to buffer short-lived context, the hippocampal system to preserve episodic memory of past experience, and internal models to imagine possible future state evolution. Inspired by these mechanisms, the authors propose MemoryVLA++, a full temporal modeling framework that equips VLA models with memory and imagination capabilities.

How It Works

1. Working memory: A pretrained VLM encodes the current observation into perceptual and cognitive tokens. 2. Episodic memory: These tokens query a perception-cognition memory bank that stores low-level details and high-level semantics of past interactions, updated through redundancy-aware consolidation. 3. Imagination: A world model imagines future states in a denoising latent space; the imagined latent variables are integrated into complete temporally aware tokens under memory guidance. 4. Action generation: These tokens condition a diffusion action expert to predict temporally consistent action sequences.

Results

Across 5 simulation benchmarks and 3 categories of real-robot tasks, the method achieves gains of:

  • +9% on general manipulation
  • +26% on memory-dependent tasks
  • +28% on imagination-dependent tasks
  • Links

  • arXiv: https://arxiv.org/abs/2506.04876

Tags

#vla#robot-manipulation#temporal-modeling#world-model#memory#paper#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981038