English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Real-Time 3D Human Motion Generation

Forum topic · 小凯 · 2026-07-13

Summary

ARDY (arXiv:2507.08713) is a streaming motion generation framework from NVIDIA Research that produces realistic, controllable 3D human motion in real time for interactive applications such as animation, simulation, and humanoid robotics. Offline text-conditioned motion generators offer precise control but run too slowly for interactivity, while existing online methods achieve real-time synthesis at the cost of controllability and long-horizon reasoning. ARDY bridges this gap using a hybrid representation combining explicit root features with latent body embeddings, balancing precise trajectory control with efficient generative learning. A two-stage autoregressive transformer denoiser with variable historical context supports conditioning on flexible, long-horizon kinematic constraints. Trained on large-scale motion capture datasets with text labels and kinematic constraints sampled from real poses, ARDY learns controllable generation natively. Evaluations on the HumanML3D benchmark and the high-fidelity Bones Rigplay dataset demonstrate high motion quality and constraint adherence. An interactive demo shows dynamic text control, keyframe pose constraints, path following, and mouse/keyboard motion control. Project page: https://research.nvidia.com/labs/sil/projects/ardy/

Overview

Field: Computer Vision Authors: Kaifeng Zhao, Mathis Petrovich, Haotian Zhang Released: 2025-07-12 arXiv: 2507.08713

Abstract

Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context windows.

Key Contributions

  • ARDY, a streaming generation framework that enables high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints.
  • A hybrid representation combining explicit root features with latent body embeddings, balancing precise trajectory control with efficient generative learning.
  • A two-stage autoregressive transformer denoiser with variable historical context, supporting conditioning on flexible, long-horizon kinematic constraints.
  • By training on large-scale motion capture datasets conditioned directly on text labels and kinematic constraints sampled from real poses, ARDY naturally learns controllable generation supporting online prompting and flexible long-term goals.
  • Results

    Extensive evaluations on the HumanML3D benchmark and the large-scale, high-fidelity Bones Rigplay dataset demonstrate ARDY's high motion quality and constraint adherence, validating key architectural decisions.

    An interactive demo showcases practical versatility, including:

  • Dynamic text control
  • Diverse keyframe pose constraints
  • Path following
  • Interactive motion control via mouse and keyboard
Supplementary videos, code, and models: https://research.nvidia.com/labs/sil/projects/ardy/

Tags

#motion-generation#diffusion-models#autoregressive#3d-human-motion#computer-vision#real-time-synthesis#nvidia#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178379420