A team from Tsinghua University, Galbot Inc., Shanghai Jiao Tong University, Peking University, and Shanghai Qi Zhi Institute has introduced Humanoid-GPT, a GPT-style generative pre-training framework for humanoid robot control that enables zero-shot, real-time tracking of highly dynamic human motions including dance and martial arts.
1. The Paradigm Shift: From Trade-offs to Scaling Laws
Humanoid whole-body control has long been stuck in a dilemma:
- Traditional policies (AMP, ASE): excellent on narrow motion categories but fail on new ones
- Generalist models (PHC, PMP): cover more motions but perform poorly on each
- Data scaling (2M → 2B frames): Transformer improves near-linearly (slight diminishing returns at 200M→2B); MLP and TCN saturate quickly.
- Model scaling: Transformer improves steadily at both small and large data; MLP/TCN overfit on small data — larger models perform worse.
The root cause is data scarcity plus architectural bottlenecks. Human motion capture is expensive, so classic models train on a few million frames — three orders of magnitude less than GPT-1's ~2 billion text tokens. The team's hypothesis: if motion can be treated as a "language" of joint angles, velocities, and positions, NLP-style Scaling Laws should apply to robot control.
2. 2 Billion Frames of Motion Data
Training data sources:
| Dataset | Scale | Notes | |---------|-------|-------| | AMASS | ~7M frames | Largest public mocap database | | LAFAN1 | ~2M frames | Long-sequence motion benchmark | | Motion-X++ | ~20M frames | Large multimodal 3D whole-body motion | | MotionMillion | Large scale | Million-level generative motion dataset | | PHUMA | Significant | Physics-grounded humanoid locomotion | | Self-collected | Large scale | In-house real-scenario data |
Preprocessing involves: 1. Filtering out robot-irrelevant motions (sitting, swimming, stair climbing) 2. Retargeting all skeletons to the 29-DoF Unitree-G1 joint space 3. Time stretching (uniform speed-up/slow-down) for 5x data augmentation
Result: 2 billion frames of G1-retargeted motion tokens.
3. Harmonic Motion Embedding (HME)
To avoid the long-tail trap of labeled categories, the team extracts periodic spectral features from the data itself:
1. Train a Periodic Autoencoder on data partitions 2. Extract per-joint periodic amplitudes and frequencies per sequence 3. Aggregate harmonic means/stds into an HME vector 4. K-Means clustering splits the 2B frames into ~300 motion clusters (~1,000–2,000 sequences each)
HME requires no motion labels — unseen motions can be automatically matched to similar existing clusters.
4. 300 RL Experts, Then Distillation
Stage 1 — 300 PPO experts: each trained on one motion cluster, with joint position/velocity observations, per-joint PD targets as actions, and a keypoint-level reward. The asymmetric reward weights reflect physical intuition: lower-body keypoints weighted 1.5x vs 0.75x upper body; rotation error weight 2.0 vs position 1.0 and velocity 0.03.
Each expert trains under extreme domain randomization: friction 0.3–2.0, terrain height up to 0.3m, random pushes every 5–10s, joint friction 0.5–2x, CoM offset ±0.15m, torso mass -3kg to +6kg.
Stage 2 — DAgger distillation: expert trajectories are distilled into a unified GPT-style causal Transformer using SmoothL1 loss, supervised in parallel over H timesteps.
5. GPT-Style Causal Transformer Architecture
| Component | Spec | |-----------|------| | Attention | Causal (masked) temporal attention | | Input | Proprioceptive state + reference pose | | History H | 32 frames (scalable to 64) | | Output | Per-joint PD targets |
Three model sizes:
| Model | Layers | Hidden | Heads | Params | |-------|--------|--------|-------|--------| | Humanoid-GPT-S | 12 | 192 | 3 | 5.38M | | Humanoid-GPT-B | 12 | 384 | 6 | 21.37M | | Humanoid-GPT-L | 12 | 768 | 12 | 85.21M |
Causal masking means early-timestep tokens see little history, so the model naturally behaves conservatively at episode start.
6. First Systematic Scaling Law Validation in Embodied AI
7. Zero-Shot Results
On the unseen AMASS-test set (Humanoid-GPT-L vs best baselines):
| Metric | Humanoid-GPT-L | Best MLP | Best TCN | |--------|----------------|----------|----------| | Success rate | 89.35% | 88.15% | 89.05% | | MPJPE | 0.0732 | 0.0832 | 0.0738 | | MPJVE | 0.5232 | 0.5285 | 0.5262 | | MPKPE | 55.15mm | 56.82mm | 56.15mm |
Real-world zero-shot demos (no fine-tuning) include dances (Michael Jackson-style, "PokerFace," "Old Town Road"), martial arts (boxing, Chinese Kungfu), basketball, single-leg jumps, daily motions (crouching, squatting, turning), and collaborative tasks (helping move/hold boxes).
8. Real-Robot Deployment
| Item | Spec | |------|------| | Inference latency | <1.5ms per forward pass | | Hardware | Single NVIDIA RTX 4090 | | Control rate | 50Hz | | Physics engine | MuJoCo | | Robot | 29-DoF Unitree-G1 |
Total training cost: ~15,000 GPU-hours (12,000 on RTX 4090s for ~384 PPO experts; 3,000 on H100s for distillation) over 3+ days on a cluster of 240 RTX 4090s + 24 H100s.
9. Limitations and Future Work
Current limits: 29-DoF-only configuration, retargeting required for new motions, purely proprioceptive (no vision), and heavy compute demands. Future directions: vision-motor fusion, multimodal (voice-to-motion) control, online learning, and cross-robot transfer.
10. Conclusion
Humanoid-GPT demonstrates that motion data can be "pre-trained" like text, that Transformer Scaling Laws hold in the physical world, and that zero-shot generalization is achievable for robots — positioning it as an embryonic "GPT-1 moment" for embodied intelligence.
Reference: Qi, Z., et al. (2026). *Humanoid Generative Pre-Training for Zero-Shot Motion Tracking*. Tsinghua University, Galbot Inc., Shanghai Jiao Tong University, Peking University, Shanghai Qi Zhi Institute.