[论文] VioLA: Learning Generalist Humanoid Control Policies from Human Data

研究领域: ML 作者: Mert Albaba, Jens Beißwenger, Anna Manasyan, Daniel Marta, Michael J. Black, Wieland Brendel, Andreas Krause, Georg Martius, Martin Riedmiller 发布时…

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Mert Albaba, Jens Beißwenger, Anna Manasyan, Daniel Marta, Michael J. Black, Wieland Brendel, Andreas Krause, Georg Martius, Martin Riedmiller 发布时间: 2026-10-08 arXiv: 2610.12435

中文摘要

教人形机器人用全身跟随指令面临两大障碍。其一,动作空间大且强耦合:腿、手臂和手指必须协同运动,同时机器人还要保持平衡,这使得关节级动作难以学习。其二,人形机器人演示数据稀缺,因此当前的人形机器人通用策略无法直接跟随新指令,部署前需要在每个任务的遥操作演示上微调。人类演示数据量大得多,但人的运动不是机器人的指令。我们通过改变通用策略的预测目标来同时移除这两个障碍。我们提出 VioLA——一个预测身体和手部运动隐变量(而非关节指令)的通用人形机器人策略。预训练的身体和手部控制器在机器人上执行这些隐变量,其对应的运动编码器将人类和机器人运动映射到相同的隐空间。因此,一段人类录像自然地在策略的动作空间中被标注,训练演示池包含 1.406 亿帧,其中 93.2% 来自人类。结果是,VioLA 在真机上零样本跟随移动指令,无需任务特定微调,达到 100% 成功率(GR00T N1.7 和 Ψ₀ 分别为 16.7% 和 0%)。在操作任务上也达到 88.6% 成功率,同样无需任务特定微调。同一方法适用于两种 VLA 和一种世界-动作模型骨干网络。仅使用人类演示训练的通用策略即可在真机上零样本执行移动任务。代码和模型检查点将公开发布。

原文摘要

Teaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn. And humanoid demonstrations are scarce, so current humanoid generalist policies do not follow new instructions out of the box and are fine-tuned on teleoperated demonstrations of each task before deployment. Human demonstrations exist in far larger numbers, but a person's motion is not a robot command. We remove both obstacles by changing what the generalist policy predicts. We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands. A pretrained body- and hand-controller exec...


*自动采集于 2026-10-11*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens