论文概要
研究领域: CV
作者: Vivek Chavan, Pengtao Xie, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger
发布时间: 2026-09-04
arXiv: 2609.05376
中文摘要
视觉运动模仿策略在分布内视觉条件下可取得高性能,但引入视觉相似物体或容器时失败。本文将此行为作为条件视觉定位问题研究:成功控制所需的视觉目标随操作阶段变化,复杂任务中还随观测任务状态变化。使用Action Chunking with Transformers(ACT),我们系统引入颜色和形状相似度可控的干扰物体和容器,并将失败定位到抓取和放置阶段。发现干扰敏感性对视觉相似类型和操作阶段均特异。基于此诊断,评估干扰增强、阶段依赖注意力正则化和基于外观的视觉提示作为互补干预,在保持控制所需空间信息的同时改善目标选择。这些干预在仿真和物理UR3e上显著提升鲁棒性。我们进一步在预训练视觉-语言-动作策略中检查相同失败模式,观测到的医疗器械状态决定正确目的地。结果表明,即使底层操作技能完好,视觉干扰也可导致错误的物体或目的地选择,显式改善目标选择可在不同视觉运动策略学习范式中大幅恢复性能。
原文摘要
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization...
自动采集于 2026-09-09
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。