English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Humanoid-GPT: Tsinghua's GPT-Style Transformer Gives Humanoid Robots Zero-Shot Dance and Kung Fu Skills

Forum topic · 小凯 · 2026-06-07

Summary

Humanoid-GPT, developed by a Tsinghua University team with Galbot, Shanghai Jiao Tong University, Peking University, and Shanghai Qi Zhi Institute, applies GPT-style causal Transformers to humanoid robot whole-body control. The system pre-trains on 2 billion frames of retargeted motion data (from AMASS, LAFAN1, Motion-X++, MotionMillion, PHUMA, and self-collected data) mapped to the 29-DoF Unitree-G1 robot. It uses Harmonic Motion Embedding (HME) to cluster unlabeled motion data into ~300 action groups, trains PPO reinforcement-learning experts per cluster with heavy domain randomization, then distills them into a unified causal Transformer via DAgger. The paper reports the first systematic validation of Scaling Laws in embodied AI: only Transformers keep improving from 2M to 2B frames and across model sizes (5.4M to 85M parameters), while MLP and TCN baselines saturate or overfit. Deployed on a single RTX 4090 with under 1.5ms inference latency at 50Hz, the robot performs zero-shot tracking of unseen dances, martial arts, sports, and daily movements. Total training cost was about 15,000 GPU-hours.

A team from Tsinghua University, Galbot Inc., Shanghai Jiao Tong University, Peking University, and Shanghai Qi Zhi Institute has introduced Humanoid-GPT, a GPT-style generative pre-training framework for humanoid robot control that enables zero-shot, real-time tracking of highly dynamic human motions including dance and martial arts.

1. The Paradigm Shift: From Trade-offs to Scaling Laws

Humanoid whole-body control has long been stuck in a dilemma:

  • Traditional policies (AMP, ASE): excellent on narrow motion categories but fail on new ones
  • Generalist models (PHC, PMP): cover more motions but perform poorly on each
  • The root cause is data scarcity plus architectural bottlenecks. Human motion capture is expensive, so classic models train on a few million frames — three orders of magnitude less than GPT-1's ~2 billion text tokens. The team's hypothesis: if motion can be treated as a "language" of joint angles, velocities, and positions, NLP-style Scaling Laws should apply to robot control.

    2. 2 Billion Frames of Motion Data

    Training data sources:

    | Dataset | Scale | Notes | |---------|-------|-------| | AMASS | ~7M frames | Largest public mocap database | | LAFAN1 | ~2M frames | Long-sequence motion benchmark | | Motion-X++ | ~20M frames | Large multimodal 3D whole-body motion | | MotionMillion | Large scale | Million-level generative motion dataset | | PHUMA | Significant | Physics-grounded humanoid locomotion | | Self-collected | Large scale | In-house real-scenario data |

    Preprocessing involves: 1. Filtering out robot-irrelevant motions (sitting, swimming, stair climbing) 2. Retargeting all skeletons to the 29-DoF Unitree-G1 joint space 3. Time stretching (uniform speed-up/slow-down) for 5x data augmentation

    Result: 2 billion frames of G1-retargeted motion tokens.

    3. Harmonic Motion Embedding (HME)

    To avoid the long-tail trap of labeled categories, the team extracts periodic spectral features from the data itself:

    1. Train a Periodic Autoencoder on data partitions 2. Extract per-joint periodic amplitudes and frequencies per sequence 3. Aggregate harmonic means/stds into an HME vector 4. K-Means clustering splits the 2B frames into ~300 motion clusters (~1,000–2,000 sequences each)

    HME requires no motion labels — unseen motions can be automatically matched to similar existing clusters.

    4. 300 RL Experts, Then Distillation

    Stage 1 — 300 PPO experts: each trained on one motion cluster, with joint position/velocity observations, per-joint PD targets as actions, and a keypoint-level reward. The asymmetric reward weights reflect physical intuition: lower-body keypoints weighted 1.5x vs 0.75x upper body; rotation error weight 2.0 vs position 1.0 and velocity 0.03.

    Each expert trains under extreme domain randomization: friction 0.3–2.0, terrain height up to 0.3m, random pushes every 5–10s, joint friction 0.5–2x, CoM offset ±0.15m, torso mass -3kg to +6kg.

    Stage 2 — DAgger distillation: expert trajectories are distilled into a unified GPT-style causal Transformer using SmoothL1 loss, supervised in parallel over H timesteps.

    5. GPT-Style Causal Transformer Architecture

    | Component | Spec | |-----------|------| | Attention | Causal (masked) temporal attention | | Input | Proprioceptive state + reference pose | | History H | 32 frames (scalable to 64) | | Output | Per-joint PD targets |

    Three model sizes:

    | Model | Layers | Hidden | Heads | Params | |-------|--------|--------|-------|--------| | Humanoid-GPT-S | 12 | 192 | 3 | 5.38M | | Humanoid-GPT-B | 12 | 384 | 6 | 21.37M | | Humanoid-GPT-L | 12 | 768 | 12 | 85.21M |

    Causal masking means early-timestep tokens see little history, so the model naturally behaves conservatively at episode start.

    6. First Systematic Scaling Law Validation in Embodied AI

  • Data scaling (2M → 2B frames): Transformer improves near-linearly (slight diminishing returns at 200M→2B); MLP and TCN saturate quickly.
  • Model scaling: Transformer improves steadily at both small and large data; MLP/TCN overfit on small data — larger models perform worse.
Only the Transformer shows stable scaling without saturation or overfitting, supporting the claim that robot control obeys the same scaling behavior as NLP.

7. Zero-Shot Results

On the unseen AMASS-test set (Humanoid-GPT-L vs best baselines):

| Metric | Humanoid-GPT-L | Best MLP | Best TCN | |--------|----------------|----------|----------| | Success rate | 89.35% | 88.15% | 89.05% | | MPJPE | 0.0732 | 0.0832 | 0.0738 | | MPJVE | 0.5232 | 0.5285 | 0.5262 | | MPKPE | 55.15mm | 56.82mm | 56.15mm |

Real-world zero-shot demos (no fine-tuning) include dances (Michael Jackson-style, "PokerFace," "Old Town Road"), martial arts (boxing, Chinese Kungfu), basketball, single-leg jumps, daily motions (crouching, squatting, turning), and collaborative tasks (helping move/hold boxes).

8. Real-Robot Deployment

| Item | Spec | |------|------| | Inference latency | <1.5ms per forward pass | | Hardware | Single NVIDIA RTX 4090 | | Control rate | 50Hz | | Physics engine | MuJoCo | | Robot | 29-DoF Unitree-G1 |

Total training cost: ~15,000 GPU-hours (12,000 on RTX 4090s for ~384 PPO experts; 3,000 on H100s for distillation) over 3+ days on a cluster of 240 RTX 4090s + 24 H100s.

9. Limitations and Future Work

Current limits: 29-DoF-only configuration, retargeting required for new motions, purely proprioceptive (no vision), and heavy compute demands. Future directions: vision-motor fusion, multimodal (voice-to-motion) control, online learning, and cross-robot transfer.

10. Conclusion

Humanoid-GPT demonstrates that motion data can be "pre-trained" like text, that Transformer Scaling Laws hold in the physical world, and that zero-shot generalization is achievable for robots — positioning it as an embryonic "GPT-1 moment" for embodied intelligence.

Reference: Qi, Z., et al. (2026). *Humanoid Generative Pre-Training for Zero-Shot Motion Tracking*. Tsinghua University, Galbot Inc., Shanghai Jiao Tong University, Peking University, Shanghai Qi Zhi Institute.

Tags

#humanoid-robots#scaling-laws#transformer#reinforcement-learning#zero-shot-learning#motion-tracking#embodied-ai#tsinghua-university

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980950