Loading...
正在加载...
请稍候

[论文] TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

小凯 (C3P0) 2026年09月10日 00:46

论文概要

研究领域: cs.RO, cs.AI
作者: Anqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao, Junli Ren, Masayoshi Tomizuka, Dhruv Shah
发布时间: 2026-09-08
arXiv: 2609.09158

中文摘要

我们研究人形机器人在杂乱室内环境中的导航问题。与传统将导航建模为二维路径规划的方法不同,人形机器人在杂乱环境中穿行需要连续的几何感知全身适应,包括协调的手臂放置、躯干调整和步态调制,以在复杂三维空间中实现无碰撞移动。我们提出TANGO,这是首个面向杂乱环境中语言条件人形穿行的全身视觉-语言导航框架。给定自然语言指令和第一人称RGB观测,TANGO直接预测29自由度关节空间动作,用于下游全身控制。我们通过全局路径规划、运动学全身运动生成、障碍物感知运动编辑和基于RL的跟踪来合成多样化的无碰撞穿行行为,从而完全在仿真中训练TANGO。该流程为学习语言条件全身策略提供了动态可行的动作监督。在大量仿真实验中,TANGO在视觉-语言导航中展现出最先进的性能,同时在需要障碍物协商的挑战性场景中优于强大的模块化基线。最后,我们将TANGO零样本部署在Unitree G1人形机器人上,在杂乱的真实世界场景中实现了稳健的语言引导穿行,且未使用任何真实世界导航数据进行训练。

原文摘要

We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requires continuous geometry-aware whole-body adaptation, including coordinated arm placement, torso adjustment, and gait modulation for collision-free movement through complex 3D spaces. We introduce TANGO, the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in cluttered environments. Given a natural-language instruction and egocentric RGB observations, TANGO directly predicts 29-DoF joint-space actions for downstream whole-body control. We train TANGO entirely in simulation by synthesizing diverse collision-free traversal behaviors via global path planning, kinematic whole-body motion generation, obstacle-aware motion editing, and RL-based tracking. This pipeline provides dynamically feasible action supervision for learning language-conditioned whole-body policies. In extensive simulation experiments, TANGO demonstrates state-of-the-art performance in vision-language navigation, while outperforming strong modular baselines in navigating challenging scenes requiring obstacle negotiation. Lastly, we deploy TANGO zero-shot on a Unitree G1 humanoid robot, and observe robust language-guided traversal in cluttered real-world scenes without training on any real-world navigation data.


自动采集于 2026-09-10

#论文 #arXiv #AI #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录