Overview
Field: Computer Vision Authors: Kaifeng Zhao, Mathis Petrovich, Haotian Zhang Released: 2025-07-12 arXiv: 2507.08713
Abstract
Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context windows.
Key Contributions
- ARDY, a streaming generation framework that enables high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints.
- A hybrid representation combining explicit root features with latent body embeddings, balancing precise trajectory control with efficient generative learning.
- A two-stage autoregressive transformer denoiser with variable historical context, supporting conditioning on flexible, long-horizon kinematic constraints.
- By training on large-scale motion capture datasets conditioned directly on text labels and kinematic constraints sampled from real poses, ARDY naturally learns controllable generation supporting online prompting and flexible long-term goals.
- Dynamic text control
- Diverse keyframe pose constraints
- Path following
- Interactive motion control via mouse and keyboard
Results
Extensive evaluations on the HumanML3D benchmark and the large-scale, high-fidelity Bones Rigplay dataset demonstrate ARDY's high motion quality and constraint adherence, validating key architectural decisions.
An interactive demo showcases practical versatility, including: