Paper Overview
Research area: Computer Vision Authors: Ilona Demler, Xinran Xie, Blake Werner Published: 2026-06-20 arXiv: 2506.16207
Abstract
The Caltech Tennis dataset (CalTennis) is a large-scale video benchmark for evaluating monocular-to-3D pose estimation in the wild. CalTennis contains over 11 million frames (51 hours) of tennis practice and matches from 40 players, captured with 2-6 synchronized cameras at 60Hz.
Key facts:
- 10x larger than existing in-the-wild human motion video datasets, and 3x larger than existing MOCAP ground-truth datasets
- The first large-scale benchmark to provide synchronized multi-view recordings of expert athletes
- The multi-view setup enables low-cost, label-free evaluation of monocular-to-3D pose estimation algorithms
- The authors present a simple, standardized data collection protocol requiring no specialized equipment or expertise, along with fully automated video calibration and synchronization
Findings
Benchmarking state-of-the-art monocular-to-3D pose estimation methods on CalTennis shows that while 3D joint angle recovery is now quite accurate, all models struggle to consistently estimate depth and foot contact. The authors further propose two new performance metrics — footwork and stability — along with a qualitative study of body-shape inconsistencies. These metrics reveal previously unexplored failure modes and point to concrete opportunities for improvement in pose estimation and motion analysis.
---
*Automatically collected on 2026-06-21*