English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

Forum topic · 小凯 · 2026-09-09

Summary

This post summarizes an arXiv paper (2609.05376) on conditional visual grounding failures in visuomotor imitation learning. Visuomotor policies such as Action Chunking with Transformers (ACT) achieve high in-distribution performance but fail when visually similar objects or receptacles are introduced. The authors systematically vary distractor color and shape similarity, localizing failures to picking and placement phases, and show that distractor sensitivity is specific to both visual similarity type and manipulation stage. Guided by this diagnosis, they evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions that improve target selection while preserving spatial control information, significantly improving robustness in simulation and on a physical UR3e robot. Similar failure modes are observed in pretrained vision-language-action policies, where the correct destination depends on observed medical device state. The results show visual distractors cause wrong object or destination selection even when manipulation skills are intact, and explicitly improving grounding largely restores performance.

Paper Overview

Field: Computer Vision Authors: Vivek Chavan, Pengtao Xie, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger Released: 2026-09-04 arXiv: 2609.05376

Summary

Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. This paper studies the behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state.

Using Action Chunking with Transformers (ACT), the authors systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to the picking and placement phases. They find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage.

Guided by this diagnosis, they evaluate three complementary interventions:

  • Distractor augmentation
  • Phase-dependent attention regularization
  • Appearance-based visual prompting
These interventions improve target selection while preserving the spatial information required for control, significantly improving robustness in simulation and on a physical UR3e robot.

The authors further examine the same failure modes in pretrained vision-language-action (VLA) policies, where the observed state of a medical device determines the correct destination. The results demonstrate that visual distractors can cause incorrect object or destination selection even when the underlying manipulation skills are intact — and that explicitly improving target selection substantially restores performance across different visuomotor policy learning paradigms.

--- *Auto-collected on 2026-09-09*

Tags

#computer-vision#visuomotor-policy#imitation-learning#visual-grounding#act#robotics#distractors#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634655