English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Large-scale Codec Avatars: High-Fidelity Full-Body 3D Avatars via Pre/Post-Training

Forum topic · 小凯 · 2026-04-04

Summary

Large-Scale Codec Avatars (LCA) is a high-fidelity, full-body 3D avatar model from researchers including Junxuan Li, Rawal Khirodkar, and Chengan He, addressing the trade-off between fidelity and generalization in avatar modeling. Multi-view studio data yields high-fidelity avatars but fails to generalize to in-the-wild inputs, while large-scale models trained on millions of wild samples generalize across identities but produce low-quality results due to 3D ambiguity. LCA introduces a large-scale pre-training and post-training paradigm for 3D avatars: pre-training on one million in-the-wild videos to learn broad appearance and geometry priors, followed by post-training on curated high-quality data to boost expressiveness and fidelity. The feedforward model generalizes across hairstyles, clothing, and demographics, offering fine-grained facial expression and finger-level articulation control with strong identity preservation. It also shows emergent generalization to relighting, loose clothing, and stylized images without direct supervision. Paper: arXiv 2504.01257.

Overview

  • Research area: Computer Vision (CV)
  • Authors: Junxuan Li, Rawal Khirodkar, Chengan He
  • Published: 2025-04-01
  • arXiv: 2504.01257

Abstract (translated)

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. Multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and the domain gap between the studio environment and the real world. On the other hand, recent large-scale avatar models trained on millions of in-the-wild samples show promise for generalization across a wide range of identities, yet the resulting avatars are often of low quality due to inherent 3D ambiguities.

To address this, the authors present Large-Scale Codec Avatars (LCA), a high-fidelity, full-body 3D avatar model that generalizes to world-scale populations in a feedforward manner, enabling efficient inference. Inspired by the success of large language models and vision foundation models, they propose, for the first time, a large-scale pre-/post-training paradigm for 3D avatar modeling: pre-training on 1 million in-the-wild videos to learn broad priors of appearance and geometry, then post-training on high-quality curated data to enhance expressiveness and fidelity.

LCA generalizes across hairstyles, clothing, and demographics, while providing precise, fine-grained facial expression and finger-level articulation control with strong identity preservation. Notably, the authors observe emergent generalization to relighting and loose clothing on unconstrained inputs, as well as zero-shot robustness to stylized images, despite the absence of direct supervision for these capabilities.

---

*Auto-collected on 2026-04-04.*

Tags

#3d-avatars#computer-vision#codec-avatars#pre-training#foundation-models#arxiv#generative-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169527