[论文] EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Camer...
研究领域: ML 作者: Kush Hari, Justin Kerr, Nidhya Shivakumar, Samarth Mahapatra, Carmelo Sferrazza, Jiahui Lei, Jitendra Malik, C. Karen Liu, Ken Goldberg, Angjoo Ka…
论文概要
研究领域: ML 作者: Kush Hari, Justin Kerr, Nidhya Shivakumar, Samarth Mahapatra, Carmelo Sferrazza, Jiahui Lei, Jitendra Malik, C. Karen Liu, Ken Goldberg, Angjoo Kanazawa 发布时间: 2026-10-02 arXiv: 2610.03710
中文摘要
受人类视觉启发,我们引入了一个利用主动凝视的框架,仅凭单目立体相机即可实现精细的双手操作。EyeRobot 2.0 通过旋转两个'眼睛'视角将其注视中心对准场景中的三维注视点,从而实现物理上的视觉注意。得到的图像被进行中心凹式处理——在图像中心分配更多视觉 token,将计算聚焦于任务相关特征。这种主动视觉注视(AVF)需要在任务执行期间精心协调凝视,我们通过分层方式实现:首先训练以目标物体为条件的低级注视伺服策略,然后训练根据任务进度发出注视目标的目标选择器。两个模块均在真实世界数据上用强化学习训练:前者使用稠密几何奖励,后者与 BC 夹爪策略共同训练,使其能发现类似人类执行任务时的注视序列。EyeRobot 2.0 进一步利用注视信息,将夹爪信息规范化到注视相对的 SE(3) 坐标系中,压缩了需要学习的动作分布规模。我们为 7 个真实任务和 6 个仿真任务收集了遥操作数据,进行了超过 1000 次物理实验和 1800 次仿真试验。移除腕部相机对标准策略代价巨大:仅有被动立体视觉时,真实世界成功率从 52% 降至 27%。EyeRobot 2.0 仅用立体相机就弥补了这一差距,在真实环境比被动立体视觉高 40%,在仿真高 20%。当腕部视野清晰时与 ego+腕部策略持平(69% vs 64%),而当抓取物体遮挡腕部相机时成功率翻倍有余(48% vs 22%)。
原文摘要
Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on real-world data: the first is trained with a dens...
*自动采集于 2026-10-06*
#论文 #arXiv #ML #小凯