English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Models

Forum topic · 小凯 · 2026-09-09

Summary

This arXiv paper (2609.05369) by Vivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger, and Jörg Krüger addresses the brittleness of vision-language-action (VLA) models in long-horizon manipulation tasks. The authors propose a neuro-symbolic framework combining learned VLA control with explicit task graphs and multimodal procedural memory. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory tracks the active step, completed actions, textual context, and task-relevant visual evidence to guide object selection, destination grounding, subgoal dispatch, and state-transition verification. Human demonstrations supply spatial-temporal guidance via gaze or saliency cues; in the initial study, pseudo-gaze is annotated directly on robot-view teleoperation videos to isolate effects on policy learning. The framework is evaluated on two long-horizon domains—workspace cleanup and surgical instrument handling—measuring correct object/destination selection, subtask completion, task progress, step-order consistency, overall success, and procedural errors.

Paper Overview

Field: Computer Vision (CV) Authors: Vivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger Published: 2026-09-04 arXiv: 2609.05369

Abstract

Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. This paper investigates a neuro-symbolic framework that combines learned VLA control with explicit task graphs and multimodal procedural memory.

Task graphs encode action dependencies, valid transitions, and branch conditions, while memory maintains the active step, completed actions, textual context, and task-relevant visual evidence. Together, these structures guide object selection, destination grounding, subgoal dispatch, and verification of expected state transitions.

Human demonstrations provide additional spatial and temporal guidance through gaze or saliency cues. To isolate the impact on policy learning, the initial study bypasses cross-view gaze transfer and annotates pseudo-gaze directly on robot-view teleoperation videos. This guidance is used for VLA fine-tuning and inference.

Evaluation Domains

The framework is studied in two long-horizon manipulation domains requiring ordered execution, visual grounding decisions, and conditional branching:

  • Workspace cleanup
  • Surgical instrument handling
  • Evaluation metrics include:

  • Correct object and destination selection
  • Subtask completion
  • Task progress
  • Step-order consistency
  • Overall task success
  • Procedural or execution errors

Conclusion

The work positions structured symbolic reasoning and demonstration-derived visual guidance as complementary mechanisms for reliable long-horizon VLA operation.

--- *Auto-collected on 2026-09-09*

Tags

#vla#neuro-symbolic#robotics#procedural-reasoning#manipulation#task-graphs#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634658