English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Latent Dynamics for Full Body Avatar Animation: Pose-Conditioned 3D Gaussian Avatars with Temporal Latent Modeling

Forum topic · 小凯 · 2026-05-22

Summary

A 2025 arXiv paper (2505.15980) by Shichong Peng, Chengxiang Yin, and Fei Jiang introduces a latent dynamics approach for pose-driven full-body avatar animation. Pose-driven neural-rendering avatars struggle with loose clothing, whose motion depends on history, inertia, and contact rather than pose alone. The proposed method augments a pose-conditioned 3D Gaussian avatar with a dynamic residual latent variable, decoded via a transformer-based network that captures temporal appearance and geometry changes beyond the driving signal. At inference, a learned latent dynamics model evolves the residual latent from a short pose history and the previous latent state, decomposing each update into driving, restoring, and dissipating forces. This yields temporally consistent, history-dependent motion at negligible extra cost, enables diverse yet plausible trajectories from different initial conditions, and exposes interpretable controls such as stiffness. Quantitative metrics and perceptual user studies on nine motion-capture sequences featuring everyday loose-clothing movements show animation quality superior to recent data-driven baselines, without requiring garment templates or test-time physics simulators.

Paper Overview

Field: Computer Vision (CV) Authors: Shichong Peng, Chengxiang Yin, Fei Jiang Published: 2025-05-20 arXiv: 2505.15980

Abstract

Pose-driven full-body avatars built on neural rendering produce high-quality novel views of a captured subject. Yet loose clothing and other dynamic elements deform in ways pose alone cannot explain: the same pose can correspond to many different states, because their motion depends on history, inertia, and contact.

Explicit simulation and layered-garment methods can model such dynamics, but they require either a dedicated garment template, which raw multi-view capture does not naturally provide, or a test-time physics simulator with non-trivial runtime cost. A parallel line of work learns data-driven clothing avatars that avoid explicit garment layers. These methods add an auxiliary latent for variation beyond pose; at inference, they fix it, regress it from pose, or retrieve it from training data, without explicitly modeling how the latent evolves with its own dynamics. Moreover, existing architectures often struggle to capture fine-grained detail in everyday loose-clothing motion, producing blurry renderings and temporal artifacts.

Method

The authors augment a pose-conditioned 3D Gaussian avatar with:

  • A transformer-based decoder that captures temporal appearance and geometry changes beyond the driving signal
  • A dynamic residual latent variable evolved at inference by a learned latent dynamics model, conditioned on a short pose history and the previous latent state
  • Each latent update is decomposed into driving, restoring, and dissipating forces, producing temporally consistent, history-dependent rollouts at negligible additional cost.

    Key Properties

  • Different initial conditions yield diverse but plausible motion trajectories
  • The force decomposition exposes interpretable controls, such as stiffness
  • No garment templates or test-time physics simulators required

Results

On nine diverse motion-capture sequences of everyday movements in loose clothing, both quantitative metrics and perceptual user studies show animation quality superior to recent data-driven baselines.

--- *Auto-collected on 2026-05-22*

Tags

#computer-vision#3d-gaussian-splatting#avatar-animation#neural-rendering#latent-dynamics#loose-clothing#motion-capture#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620577