English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ETCH-X: Robust Expressive Body Fitting for Clothed Humans with Composable Datasets

Forum topic · 小凯 · 2026-04-11

Summary

ETCH-X is an upgraded human body fitting method that aligns the expressive SMPL-X parametric model to raw 3D point clouds of clothed humans. Building on ETCH, it introduces a tightness-aware fitting paradigm to filter out clothing dynamics ("undress"), and replaces explicit sparse markers, which are highly sensitive to partial data, with implicit dense correspondences ("dense fit"). Its disentangled modular stages allow separate, scalable training on composable data sources: simulated garments (CLOTH3D), large-scale full-body motions (AMASS), and fine-grained hand gestures (InterHand2.6M), improving outfit generalization and pose robustness for both bodies and hands. Compared with ETCH, ETCH-X substantially reduces error on seen data such as 4D-Dress (MPJPE-All down 33.0%) and CAPE (V2V-Hands down 35.8%), and on unseen data such as BEDLAM2.0 (MPJPE-All down 80.8%; V2V-All down 80.5%), achieving robust, expressive fitting across diverse clothing, poses, and input completeness levels. Paper: arXiv 2504.07086.

Paper Overview

  • Field: AI / Computer Vision
  • Authors: Xiaoben Li, Jingyi Wu, Zeyu Cai
  • Published: 2025-04-10
  • arXiv: 2504.07086
  • Abstract

    Human body fitting, which aligns parametric body models such as SMPL to raw 3D point clouds of clothed humans, serves as a crucial first step for downstream tasks like animation and texturing. An effective fitting method should be both locally expressive — capturing fine details such as hands and facial features — and globally robust to handle real-world challenges, including clothing dynamics, pose variations, and noisy or partial inputs. Existing approaches typically excel in only one aspect, lacking an all-in-one solution.

    Key Contributions

    The authors upgrade ETCH to ETCH-X, which:

  • Leverages a tightness-aware fitting paradigm to filter out clothing dynamics ("undress").
  • Extends expressiveness with SMPL-X.
  • Replaces explicit sparse markers, which are highly sensitive to partial data, with implicit dense correspondences ("dense fit") for more robust and fine-grained body fitting.
  • Uses disentangled "undress" and "dense fit" modular stages that enable separate and scalable training on composable data sources: diverse simulated garments (CLOTH3D), large-scale full-body motions (AMASS), and fine-grained hand gestures (InterHand2.6M), improving outfit generalization and pose robustness of both bodies and hands.
  • Results

    ETCH-X achieves robust and expressive fitting across diverse clothing, poses, and levels of input completeness, delivering substantial improvements over ETCH on:

    1. Seen data: 4D-Dress (MPJPE-All reduced by 33.0%) and CAPE (V2V-Hands reduced by 35.8%). 2. Unseen data: BEDLAM2.0 (MPJPE-All reduced by 80.8%; V2V-All reduced by 80.5%).

    Links

  • Paper: https://arxiv.org/abs/2504.07086

Tags

#computer-vision#3d-human-body-fitting#smpl-x#point-clouds#arxiv#papers#cloth3d#amass

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169732