Overview
- Research area: Computer Vision (CV)
- Authors: Junxuan Li, Rawal Khirodkar, Chengan He
- Published: 2025-04-01
- arXiv: 2504.01257
Abstract (translated)
High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. Multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and the domain gap between the studio environment and the real world. On the other hand, recent large-scale avatar models trained on millions of in-the-wild samples show promise for generalization across a wide range of identities, yet the resulting avatars are often of low quality due to inherent 3D ambiguities.
To address this, the authors present Large-Scale Codec Avatars (LCA), a high-fidelity, full-body 3D avatar model that generalizes to world-scale populations in a feedforward manner, enabling efficient inference. Inspired by the success of large language models and vision foundation models, they propose, for the first time, a large-scale pre-/post-training paradigm for 3D avatar modeling: pre-training on 1 million in-the-wild videos to learn broad priors of appearance and geometry, then post-training on high-quality curated data to enhance expressiveness and fidelity.
LCA generalizes across hairstyles, clothing, and demographics, while providing precise, fine-grained facial expression and finger-level articulation control with strong identity preservation. Notably, the authors observe emergent generalization to relighting and loose clothing on unconstrained inputs, as well as zero-shot robustness to stylized images, despite the absence of direct supervision for these capabilities.
---
*Auto-collected on 2026-04-04.*