English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PRIME: Perception Feedback with Situational Memory Embeddings for Vision-Language-Action Autonomous Driving Models

Forum topic · 小凯 · 2026-09-22

Summary

PRIME is a learned feedback mechanism for vision-language-action (VLA) models in autonomous driving, introduced by researchers including Erik Deinzer and Luigi Palmieri (arXiv 2609.22040). Current VLA models run feedforward inference across a perception-reasoning-planning hierarchy: early perception processes visual inputs without awareness of downstream reasoning or navigation goals. PRIME bridges this gap by conditioning perceptual queries on a novel Situational Memory, aggregating latent representations of past perception, reasoning, navigation goals, and predicted behaviors over an L-step window via cross-attention. This enables intent-driven perceptual attention while adding at most 29.7M parameters—just 0.41% of the 7.3B-parameter base model. On the Bench2Drive closed-loop benchmark, PRIME achieves a state-of-the-art driving score of 82.47 (4.73 above ORION) and a 60.00% success rate (5.38 points higher), the highest reported driving score among published VLAs trained on Think2Drive demonstrations.

Overview

Field: Computer Vision (CV) Authors: Erik Deinzer, Naya Baslan, Luca Paparusso, Narunas Vaskevicius, Peter Knott, Luigi Palmieri arXiv: 2609.22040

Background

Current Vision-Language-Action (VLA) models for autonomous driving operate primarily through feedforward inference across the perception–reasoning–planning hierarchy. While modern architectures maintain temporal recurrence within the perceptual module, early perception remains blind to downstream reasoning and navigation goals, processing visual inputs agnostically without prioritizing cues informed by prior decisions.

Method: PRIME

PRIME introduces a learned feedback mechanism that conditions the VLA perceptual queries on a novel Situational Memory:

  • Aggregates latent representations of past perception, reasoning, navigation goals, and predicted behaviors across an L-step window via cross-attention.
  • Enables intent-driven perceptual attention at minimal computational cost.
  • Adds at most 29.7M parameters — only 0.41% of the 7.3B-parameter base model.
  • Results

    Evaluated on the Bench2Drive closed-loop benchmark:

  • Driving score: 82.47 (state-of-the-art, +4.73 over ORION)
  • Success rate: 60.00% (+5.38 points)
  • Highest reported driving score among all published VLAs trained on Think2Drive demonstrations.
  • Links

  • arXiv: https://arxiv.org/abs/2609.22040
*Auto-collected on 2026-09-22.*

Tags

#autonomous-driving#vision-language-action#situational-memory#perception-feedback#bench2drive#arxiv#machine-learning#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635076