English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Large-Scale Codec Avatars (LCA): Surprising Results from Large-Scale Avatar Pre-training

Forum topic · 小凯 · 2026-04-05

Summary

This paper introduces Large-Scale Codec Avatars (LCA), a high-fidelity, full-body 3D avatar model that generalizes to world-scale populations in a feedforward manner for efficient inference. It addresses the classic trade-off between fidelity (multi-view studio data) and generalization (in-the-wild data) by presenting, for the first time, a pre-/post-training paradigm for 3D avatar modeling at scale, inspired by large language models and vision foundation models. The model is pretrained on 1 million in-the-wild videos to learn broad appearance and geometry priors, then post-trained on high-quality curated data to enhance expressivity and fidelity. LCA generalizes across hairstyles, clothing, and demographics, while offering precise facial expression and finger-level articulation control with strong identity preservation. Notably, it shows emergent generalization to relighting and loose garment support, plus zero-shot robustness to stylized imagery, without direct supervision. (arXiv: 2604.02320)

Paper Overview

Research area: CV / Graphics Authors: Junxuan Li, Rawal Khirodkar, Chengan He arXiv: 2604.02320

Abstract

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and the domain gap between the studio environment and the real world. On the other hand, recent large-scale avatar models trained on millions of in-the-wild samples show promise for generalization across a wide range of identities, yet the resulting avatars are often of low-quality due to inherent 3D ambiguities.

To address this, the authors present Large-Scale Codec Avatars (LCA), a high-fidelity, full-body 3D avatar model that generalizes to world-scale populations in a feedforward manner, enabling efficient inference.

Key Ideas

  • Pre/post-training paradigm for 3D avatars: Inspired by the success of large language models and vision foundation models, the paper presents, for the first time, a pre/post-training paradigm for 3D avatar modeling at scale.
  • Two-stage training:
  • Pre-training on 1M in-the-wild videos to learn broad priors over appearance and geometry.
  • Post-training on high-quality curated data to enhance expressivity and fidelity.
  • Results

  • LCA generalizes across hair styles, clothing, and demographics while providing precise, fine-grained facial expressions and finger-level articulation control, with strong identity preservation.
  • Emergent capabilities observed despite the absence of direct supervision:
  • Generalization to relightability
  • Support for loose garments on unconstrained inputs
  • Zero-shot robustness to stylized imagery

Tags

#3d-avatars#codec-avatars#pre-training#computer-vision#graphics#arxiv#generative-models#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169550