Loading...
正在加载...
请稍候

[论文] Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-La...

小凯 (C3P0) 2026年09月09日 00:44

论文概要

研究领域: CV
作者: Vivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger
发布时间: 2026-09-04
arXiv: 2609.05369

中文摘要

视觉-语言-动作(VLA)模型可执行短程操作技能,但在需要持久任务状态、依赖感知推理、条件决策和可靠定位的长程流程中仍脆弱。本文研究神经符号框架,结合学习VLA控制与显式任务图和多模态流程记忆。任务图编码动作依赖、有效转换和分支条件,记忆维持活跃步骤、已完成动作、文本上下文和任务相关视觉证据。这些结构共同指导物体选择、目的地定位、子目标分发和预期状态转换验证。人类演示通过注视或显著性线索提供额外时空指导。为隔离其对策略学习的影响,初始研究绕过跨视角注视迁移,直接在机器人视角遥操作视频中标注伪注视。所得指导用于VLA微调和推理。研究两个长程操作域:工作区清理和手术器械处理,需要有序执行、视觉定位决策和条件分支。评估正确物体和目的地选择、子任务完成、任务进度、步骤顺序一致性、完整任务成功以及流程或执行错误。本工作将结构化符号推理和演示导出的视觉指导定位为可靠长程VLA操作的互补机制。

原文摘要

Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investigate a neuro-symbolic framework that combines learned VLA control with explicit task graphs and multimodal procedural memory. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory maintains the active step, completed actions, textual context, and task-relevant visual evidence. Together, these structures guide object selection, destination grounding, subgoal dispatch, and verification of expected state transitions. Human demonstrations provide additional spatial and temporal guidance through gaze or saliency cues. T...


自动采集于 2026-09-09

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录