Loading...
正在加载...
请稍候

[论文] HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation...

小凯 (C3P0) 2026年08月19日 00:56

论文概要

研究领域: Robotics
作者: Langzhe Gu, Chengkai Hou, Meng Li et al. (17 authors)
发布时间: 2026-08-17
arXiv: 2608.16837

中文摘要

人形机器人在以人为中心的环境中作为通用智能体大有前途,但通用视觉-语言-动作(VLA)基础模型并不直接适用于人形全身运动-操作。人形运动的高维度和相互依赖性使得传统单阶段VLA架构难以有效协调运动、腰部姿态和双臂操作。此外,通过离线行为克隆训练的策略在真实世界部署期间可能保持次优。虽然在线强化学习可通过真实世界交互细化策略,但直接微调大型VLA骨干需要过多计算,并可能在真实机器人探索期间引入安全风险。为解决这些瓶颈,我们引入HAF(人形适应框架),由HAF-VLA和HAF-Steer两部分组成,将现成通用VLA基础模型迁移到人形全身运动-操作。HAF-VLA是基于预训练流匹配VLA的层次化动作流生成器,将全身动作去噪分为三个阶段,保留运动学依赖,避免一次性生成不连贯的全身动作。在冻结的HAF-VLA之上,HAF-Steer是一个潜离线到在线RL管道,利用流匹配可逆性和基于DCT的降维将RL优化限制在紧凑噪声子空间。在七个真实世界人形运动-操作任务上评估,HAF超越了普通单阶段VLA基线。

原文摘要

Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively. Moreover, policies trained through offline behavior cloning can remain suboptimal during real-world deployment. Although online reinforcement learning can refine policies through real-world interaction, directly tuning large VLA backbones demands excessive computation and may introduce safety risks during real-robot exploration. To address these bottlenecks, we ...


自动采集于 2026-08-19

#论文 #arXiv #Robotics #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录