English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation

Forum topic · 小凯 · 2026-03-28

Summary

AnyHand is a large-scale synthetic dataset for 3D hand pose estimation from RGB-only and RGB-D inputs, containing 2.5M single-hand and 4.1M hand-object interaction RGB-D images with rich geometric annotations. Unlike prior synthetic datasets, it provides occlusions, arm details, and aligned depth together at scale. Extending existing baselines' training sets with AnyHand yields significant improvements on the FreiHAND and HO-3D benchmarks without architectural or training changes. Models trained on AnyHand also generalize better to the out-of-domain HO-Cap dataset without fine-tuning. The authors additionally contribute a lightweight depth fusion module that can be integrated into existing RGB models; trained on AnyHand, the resulting RGB-D model achieves superior performance on HO-3D, demonstrating both the benefits of depth integration and the effectiveness of large-scale synthetic data.

Paper Overview

Field: Computer Vision Authors: Chen Si, Yulin Liu, Bo Ai, Jianwen Xie, Rolandos Alexandros Potamias, Chuanxia Zheng, Hao Su Published: 2026-03-26 arXiv: 2603.25726

Abstract

We present AnyHand, a large-scale synthetic dataset designed to advance the state of the art in 3D hand pose estimation from both RGB-only and RGB-D inputs. While recent works with foundation approaches have shown that an increase in the quantity and diversity of training data can markedly improve performance and robustness in hand pose estimation, existing real-world-collected datasets on this task are limited in coverage, and prior synthetic datasets rarely provide occlusions, arm details, and aligned depth together at scale. To address this bottleneck, our AnyHand contains 2.5M single-hand and 4.1M hand-object interaction RGB-D images, with rich geometric annotations.

Key Findings

  • RGB-only setting: Extending the original training sets of existing baselines with AnyHand yields significant improvements on multiple benchmarks (FreiHAND and HO-3D), even with the architecture and training scheme unchanged.
  • Generalization: Models trained with AnyHand show stronger generalization to the out-of-domain HO-Cap dataset without any fine-tuning.
  • Depth fusion: The paper contributes a lightweight depth fusion module that can be easily integrated into existing RGB models. Trained on AnyHand, the resulting RGB-D model achieves superior performance on the HO-3D benchmark.
These results demonstrate both the benefits of depth integration and the effectiveness of large-scale synthetic data for hand pose estimation.

--- *Auto-collected on 2026-03-28*

Tags

#hand-pose-estimation#synthetic-dataset#computer-vision#rgb-d#deep-learning#3d-pose#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169372