English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking in Humanoid Whole-Body Control

Forum topic · 小凯 · 2026-06-04

Summary

Humanoid-GPT is a GPT-style Transformer with causal attention, trained on a billion-scale motion corpus for humanoid whole-body control. Unlike prior shallow MLP-based motion trackers limited by scarce training data and an agility-generalization trade-off, Humanoid-GPT is pre-trained on a 2-billion-frame retargeted corpus that unifies all major motion capture datasets with large-scale in-house recordings. By scaling both data volume and model capacity, the authors produce a single generative Transformer capable of tracking highly dynamic behaviors while achieving strong zero-shot generalization to unseen motions and control tasks. Extensive experiments and scaling analyses indicate the model sets a new performance frontier, combining robust zero-shot transfer to unseen tasks with accurate tracking of highly dynamic and complex movements. The paper was authored by Zekun Qi, He Wang, Li Yi, and colleagues, and posted to arXiv as 2606.03985 (June 2026). This post shares the paper abstract and links for the computer vision and robotics community on zhichai.net.

Paper Overview

  • Field: Computer Vision (CV)
  • arXiv: 2606.03985
  • Authors: Zekun Qi, Xuchuan Chen, Dairu Liu, Chenghuai Lin, Yunrui Lian, Sikai Liang, Zhikai Zhang, Yu Guan, Jilong Wang, Wenyao Zhang, Xinqiang Yu, He Wang, Li Yi
  • Summary

    We introduce Humanoid-GPT, a GPT-style Transformer with causal attention trained on a billion-scale motion corpus for whole-body control. Unlike prior shallow MLP trackers constrained by scarce data and an agility-generalization trade-off, Humanoid-GPT is pre-trained on a 2B-frame retargeted corpus that unifies all major mocap datasets with large-scale in-house recordings.

    Scaling both data and model capacity yields a single generative Transformer that tracks highly dynamic behaviors while achieving unprecedented zero-shot generalization to unseen motions and control tasks. Extensive experiments and scaling analyses show that our model establishes a new performance frontier, demonstrating robust zero-shot generalization to unseen tasks while simultaneously tracking highly dynamic and complex motion.

    Key Takeaways

  • A GPT-style causal-attention Transformer applied to humanoid whole-body motion tracking
  • Pre-trained on a 2B-frame retargeted corpus spanning all major mocap datasets plus in-house recordings
  • Overcomes the agility-generalization trade-off seen in shallow MLP trackers
  • Achieves zero-shot generalization to unseen motions and control tasks while tracking highly dynamic behaviors

Tags

#humanoid-robotics#motion-tracking#whole-body-control#gpt#transformer#zero-shot-learning#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980807