English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CalTennis: Large Multi-View Tennis Video Dataset and Benchmark for Monocular-to-3D Pose Estimation

Forum topic · 小凯 · 2026-06-22

Summary

CalTennis is a large-scale video benchmark for evaluating monocular-to-3D human pose estimation in the wild, built from tennis practice and match footage. The dataset contains over 11 million frames (about 51 hours) from 40 players, captured at 60 Hz with 2-6 synchronized cameras. It is roughly 10x larger than existing in-the-wild human motion video datasets and about 3x larger than datasets with MOCAP ground-truth annotations, and is the first large-scale benchmark to provide synchronized multi-view recordings of expert athletic motion. The multi-view setup enables low-cost, markerless evaluation of monocular 3D pose algorithms, supported by a simple standardized capture protocol and fully automatic calibration and synchronization. Benchmarking state-of-the-art monocular 3D pose methods on CalTennis shows that 3D joint angle recovery has become accurate, while models remain unreliable at depth estimation and foot-ground contact. The authors propose two novel metrics, footwork and stability, and qualitatively examine body shape inconsistency, revealing previously unexplored failure modes and pointing to concrete improvements for pose estimation and motion analysis.

Paper Overview

Field: cs.CV Authors: Ilona Demler, Xinran Xie, Blake Werner Published: 2026-06-21 arXiv: 2506.17579

Summary

CalTennis is a large-scale video benchmark for evaluating monocular-to-3D pose estimation in the wild.

CalTennis compiles footage of tennis practice and matches from 40 players — over 11 million frames, totaling 51 hours — captured with 2 to 6 synchronized cameras at 60 Hz. It is more than 10x larger than existing in-the-wild human motion video datasets and roughly 3x larger than datasets with motion capture (MOCAP) ground-truth annotations. It is also the first large-scale benchmark to provide synchronized multi-view recordings of expert athletic motion.

The multi-view setup enables low-cost, markerless evaluation of monocular 3D pose estimation algorithms. The authors describe a simple, standardized capture protocol that requires no specialized equipment or expertise, combined with fully automatic video calibration and synchronization.

Benchmarking state-of-the-art monocular 3D pose methods on CalTennis, the authors find that 3D joint angle recovery has become accurate, but models still struggle with depth estimation and foot-ground contact detection. They further propose two novel performance metrics — footwork and stability — and qualitatively investigate body shape inconsistency. These metrics expose previously unexplored failure modes and point to concrete directions for improving pose estimation and motion analysis.

--- *Automatically collected on 2026-06-21*

Tags

#pose-estimation#3d-vision#computer-vision#dataset#benchmark#tennis#multi-view#motion-analysis

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178207995