Paper Overview
- Field: Computer Vision (CV)
- arXiv: 2606.03985
- Authors: Zekun Qi, Xuchuan Chen, Dairu Liu, Chenghuai Lin, Yunrui Lian, Sikai Liang, Zhikai Zhang, Yu Guan, Jilong Wang, Wenyao Zhang, Xinqiang Yu, He Wang, Li Yi
- A GPT-style causal-attention Transformer applied to humanoid whole-body motion tracking
- Pre-trained on a 2B-frame retargeted corpus spanning all major mocap datasets plus in-house recordings
- Overcomes the agility-generalization trade-off seen in shallow MLP trackers
- Achieves zero-shot generalization to unseen motions and control tasks while tracking highly dynamic behaviors
Summary
We introduce Humanoid-GPT, a GPT-style Transformer with causal attention trained on a billion-scale motion corpus for whole-body control. Unlike prior shallow MLP trackers constrained by scarce data and an agility-generalization trade-off, Humanoid-GPT is pre-trained on a 2B-frame retargeted corpus that unifies all major mocap datasets with large-scale in-house recordings.
Scaling both data and model capacity yields a single generative Transformer that tracks highly dynamic behaviors while achieving unprecedented zero-shot generalization to unseen motions and control tasks. Extensive experiments and scaling analyses show that our model establishes a new performance frontier, demonstrating robust zero-shot generalization to unseen tasks while simultaneously tracking highly dynamic and complex motion.