[论文] Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-La...

研究领域: CV 作者: Vivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger 发布时间: 2026-09-04 arXiv: 2609.05369

论文概要

研究领域: CV 作者: Vivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger 发布时间: 2026-09-04 arXiv: 2609.05369

中文摘要

视觉-语言-动作(VLA)模型可执行短程操作技能,但在需要持久任务状态、依赖感知推理、条件决策和可靠定位的长程流程中仍脆弱。本文研究神经符号框架,结合学习VLA控制与显式任务图和多模态流程记忆。任务图编码动作依赖、有效转换和分支条件,记忆维持活跃步骤、已完成动作、文本上下文和任务相关视觉证据。这些结构共同指导物体选择、目的地定位、子目标分发和预期状态转换验证。人类演示通过注视或显著性线索提供额外时空指导。为隔离其对策略学习的影响,初始研究绕过跨视角注视迁移,直接在机器人视角遥操作视频中标注伪注视。所得指导用于VLA微调和推理。研究两个长程操作域:工作区清理和手术器械处理,需要有序执行、视觉定位决策和条件分支。评估正确物体和目的地选择、子任务完成、任务进度、步骤顺序一致性、完整任务成功以及流程或执行错误。本工作将结构化符号推理和演示导出的视觉指导定位为可靠长程VLA操作的互补机制。

原文摘要

Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investigate a neuro-symbolic framework that combines learned VLA control with explicit task graphs and multimodal procedural memory. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory maintains the active step, completed actions, textual context, and task-relevant visual evidence. Together, these structures guide object selection, destination grounding, subgoal dispatch, and verification of expected state transitions. Human demonstrations provide additional spatial and temporal guidance through gaze or saliency cues. T...


*自动采集于 2026-09-09*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens