LIMMT Deep Dive: When 3% of Data Beats 100% — The 'Less is More' Paradigm for Motion Tracking
> Paper: LIMMT: Less is More for Motion Tracking > arXiv: https://arxiv.org/abs/2606.06953 > Project page: https://giraffeguan.github.io/limmt/ > Authors: Yu Guan, Zekun Qi, Chenghuai Lin, Xuchuan Chen, Dairu Liu, Wenyao Zhang, Jilong Wang, Xinqiang Yu, He Wang, Li Yi > Institutions: Tsinghua University, GalBot, Peking University, Shanghai Qi Zhi Institute, Shanghai Jiao Tong University, University of Edinburgh
Key points
- Core claim: In physics-based imitation RL for motion tracking, training signal quality matters more than data quantity. LIMMT selects ~3% of AMASS (~420 of 14,000 motions) and beats training on 100% of it.
- Results: | Data strategy | Success (Any2Track) | Success (TWIST2) | |---|---|---| | Full AMASS (100%) | 94.2% | 82.5% | | Random 3% | 83.8% | 64.9% | | GQS 3% | 95.6% | 86.1% | | GQS 10% | 95.9% | 86.8% |
- PHUMA (in-domain): GQS with 30% of data exceeds full-dataset performance (which is near ceiling at 99.31%).
- Zero-shot transfer PHUMA → AMASS: GQS 10% subset reaches 92.8% vs 91.0% for full data — the curated subset is *more* robust, since removing redundant easy data prevents overfitting to source-domain artifacts.
- Sim-to-real: Policies trained on GQS 10% data in MJX/Isaac Lab transfer zero-shot to a real Unitree G1 humanoid, covering everyday, dance, and athletic motions.
- Paper: https://arxiv.org/abs/2606.06953
- Project page: https://giraffeguan.github.io/limmt/
- AMASS: https://amass.is.tue.mpg.de/
- PHUMA: https://github.com/yusun-nlp/PHUMA
- Any2Track: https://arxiv.org/abs/2509.13833
- TWIST2: https://arxiv.org/abs/2511.02832
Random subsampling is catastrophic — the problem is bad data, not too little data.
The GQS (General Quality Selection) pipeline
LIMMT defines three dimensions of motion data value, processed in a strict order:
Stage I: Physics feasibility filtering
Not a binary filter but a soft score Sphy(T) = 100 - Σ wi·Li, penalizing physics violations detected by replaying each motion in a rigid-body simulator:
| Metric | Weight | Role | |---|---|---| | Floating | 24.19 | Highly toxic (catastrophic reconstruction error) | | Penetration | 216.62 | Neutral (slight penetration acceptable) | | Joint velocity | 44.22 | Friendly (high-speed motion is a rich signal) | | Foot slide | 1.70 | Harmful | | Jerk | 0.28 | Neutral | | Self-collision | 0.17 | Friendly |
Counter-intuitive finding: high-velocity motions are "friendly" — the goal is removing *impossible* motions, not *imperfect* ones.
Stage II: Harmonic Motion Embedding (HME)
A periodic autoencoder decomposes each joint signal as z_i(t) = A_i·sin(2π(F_i·t + φ_i)) + b_i. The global embedding uses only amplitude and frequency means (z_global = (1/N) Σ_w [A_w, F_w]), giving phase invariance — the same motion starting at different times maps close together.
Stage III: Complexity-weighted farthest point sampling
Iterative FPS with mixed score Score(u) = α·D̂(u,S) + (1-α)·Ĉ(u), where D̂ is embedding-space diversity distance and Ĉ is complexity (kinetic energy + λ·acceleration). With α = 0.99, diversity dominates; complexity breaks ties toward dynamically richer motions.
Ablations
Removing any stage in the 3% regime hurts: no physics filtering drops success to 91.1% (embedding sampling favors outliers), no diversity sparsification to 93.4%, no complexity weighting to 94.6%. All three combined: 95.6%.
A notable non-monotonic result: training on motions binned by physics score, the 60-70% bin performs best (96.3%), while the highest-quality bin (0-10% of violations) underperforms (94.6%) because such motions tend to be static and conservative, and the worst bin collapses (92.2%). Physics scores identify toxic data but cannot rank viable motions — validating the three-stage design.
Cross-domain and real-robot validation
Training-curve analysis shows GQS subsets achieve higher reward and lower tracking error from early training onward — good data changes the optimization trajectory, not just convergence speed.
Paradigm implications
LIMMT shifts from *data engineering* (collect more, clean obvious errors) to *data curation* (define value dimensions, optimize information density). The authors suggest the principle — toxic gradients mislead optimization — may extend to vision pretraining, LLM data selection, robot imitation learning, and RL trajectory curation, though the specific dimensions are physics-RL-specific.
Limitations: GQS is a static, one-shot pre-training filter; the Sphy ≥ 90 threshold is hand-tuned; curation requires simulator rollouts of every candidate; and the framework is specific to motion tracking.