Paper Overview
Field: Computer Vision Authors: Chen Si, Yulin Liu, Bo Ai, Jianwen Xie, Rolandos Alexandros Potamias, Chuanxia Zheng, Hao Su Published: 2026-03-26 arXiv: 2603.25726
Abstract
We present AnyHand, a large-scale synthetic dataset designed to advance the state of the art in 3D hand pose estimation from both RGB-only and RGB-D inputs. While recent works with foundation approaches have shown that an increase in the quantity and diversity of training data can markedly improve performance and robustness in hand pose estimation, existing real-world-collected datasets on this task are limited in coverage, and prior synthetic datasets rarely provide occlusions, arm details, and aligned depth together at scale. To address this bottleneck, our AnyHand contains 2.5M single-hand and 4.1M hand-object interaction RGB-D images, with rich geometric annotations.
Key Findings
- RGB-only setting: Extending the original training sets of existing baselines with AnyHand yields significant improvements on multiple benchmarks (FreiHAND and HO-3D), even with the architecture and training scheme unchanged.
- Generalization: Models trained with AnyHand show stronger generalization to the out-of-domain HO-Cap dataset without any fine-tuning.
- Depth fusion: The paper contributes a lightweight depth fusion module that can be easily integrated into existing RGB models. Trained on AnyHand, the resulting RGB-D model achieves superior performance on the HO-3D benchmark.
--- *Auto-collected on 2026-03-28*