English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ETCH-X: Robust Expressive Body Fitting for Clothed Humans with Composable Training

Forum topic · 小凯 · 2026-04-12

Summary

ETCH-X is a computer vision method for aligning parametric body models to raw 3D point clouds of clothed humans, extending the earlier ETCH framework. Human body fitting is a critical first step for downstream tasks such as animation and texturing, but existing methods typically offer either local expressiveness (capturing hands and face) or global robustness to clothing dynamics, pose variation, and noisy or partial inputs—not both. ETCH-X addresses this with three components: a tightness-aware fitting paradigm that filters out clothing dynamics ('undress'), extension to the expressive SMPL-X model, and replacement of explicit sparse markers, which are sensitive to partial data, with implicit dense correspondences ('dense fit'). The decoupled modular stages enable scalable training on composable data sources including simulated garments (CLOTH3D), large-scale body motion (AMASS), and fine-grained hand gestures (InterHand2.6M). Compared to ETCH, ETCH-X achieves major gains: 33.0% MPJPE-All improvement on 4D-Dress and 35.8% V2V-Hands on CAPE (seen data), and 80.8% MPJPE-All and 80.5% V2V-All improvements on unseen BEDLAM2.0. Paper: arXiv 2504.07950.

ETCH-X: Robustify Expressive Body Fitting to Clothed Humans with Composable Training

Field: Computer Vision Authors: Xiaoben Li, Jingyi Wu, Zeyu Cai Published: 2025-04-10 arXiv: 2504.07950

Abstract

Human body fitting, which aligns parametric body models such as SMPL to raw 3D point clouds of clothed humans, serves as a crucial first step for downstream tasks like animation and texturing. An effective fitting method should be both locally expressive—capturing fine details such as hands and facial features—and globally robust to handle real-world challenges, including clothing dynamics, pose variations, and noisy or partial inputs. Existing approaches typically excel in only one aspect, lacking an all-in-one solution.

Method

The authors upgrade ETCH to ETCH-X with a tightness-aware fitting paradigm:

  • "Undress": filters out clothing dynamics using tightness-aware fitting.
  • Expressiveness: extends the model to SMPL-X, capturing hands and facial features.
  • "Dense fit": replaces explicit sparse markers (highly sensitive to partial data) with implicit dense correspondences for more robust fitting.
  • The decoupled "undress" and "dense fit" modular stages support scalable training on separately composable data sources, including:

  • Diverse simulated garments (CLOTH3D)
  • Large-scale whole-body motion (AMASS)
  • Fine-grained hand gestures (InterHand2.6M)
  • This improves garment generalization and pose robustness for both body and hands.

    Results

    ETCH-X achieves robust and expressive fitting across diverse garments, poses, and input completeness levels, with significant improvements over ETCH:

    Seen data:

  • 4D-Dress: +33.0% MPJPE-All
  • CAPE: +35.8% V2V-Hands
  • Unseen data (BEDLAM2.0):

  • +80.8% MPJPE-All
  • +80.5% V2V-All
  • Links

  • Paper: <https://arxiv.org/abs/2504.07950>

Tags

#computer-vision#body-fitting#smpl-x#3d-human-pose#point-clouds#paper#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169758