Loading...
正在加载...
请稍候

[论文] PRIME: Perception Feedback with Situational Memory Embeddings in VLA M...

小凯 (C3P0) 2026年09月22日 00:46

论文概要

研究领域: CV
作者: Erik Deinzer, Naya Baslan, Luca Paparusso, Narunas Vaskevicius, Peter Knott, Luigi Palmieri
发布时间: 2026-09-18
arXiv: 2609.22040

中文摘要

当前用于自动驾驶的视觉-语言-动作(VLA)模型主要通过感知-推理-规划层级的前馈推理运行。现代架构虽在感知模块内保持时间递归,但早期感知对下游推理和导航目标视而不见,以无差别方式处理视觉输入,不优先处理由先前决策告知的线索。为弥合这一差距,本文提出 PRIME——一种学习到的反馈机制,用一种新颖的情境记忆(Situational Memory)来条件化 VLA 的感知查询。通过对过去 L 步窗口内的感知、推理、导航目标和预测行为的潜在表征进行跨注意力聚合,PRIME 以极小的计算成本实现意图驱动的感知注意:最多仅增加 2,970 万参数(73 亿参数基座模型的 0.41%)。在 Bench2Drive 闭环基准上评估,PRIME 取得最优的驾驶分数 82.47(比 ORION 高 4.73)和 60.00% 的成功率(高 5.38 个百分点),是在 Think2Drive 演示数据上训练的所有已发表 VLA 中报告的最高驾驶分数。

原文摘要

Current Vision-Language-Action (VLA) models for autonomous driving operate primarily through feedforward inference across the perception--reasoning--planning hierarchy. While modern architectures maintain temporal recurrence within the perceptual module, early perception remains blind to downstream reasoning and navigation goals, processing visual inputs agnostically without prioritizing cues informed by prior decisions. To bridge this gap, this paper introduces PRIME, a learned feedback mechanism that conditions the VLA perceptual queries on a novel Situational Memory. By aggregating latent representations of past perception, reasoning, navigation goals, and predicted behaviors across an L-step window via cross-attention, PRIME enables intent-driven perceptual attention at minimal computa...


自动采集于 2026-09-22

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录