English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CalTennis: Large Multi-View Tennis Video Dataset and Benchmark for Monocular-to-3D Pose Estimation

Forum topic · 小凯 · 2026-06-21

Summary

CalTennis is a large-scale video benchmark for evaluating monocular-to-3D human pose estimation in the wild, introduced by researchers at Caltech (arXiv:2506.16207). The dataset contains over 11 million frames (51 hours) of tennis practice and matches from 40 players, captured with 2-6 synchronized cameras at 60Hz. It is 10x larger than existing in-the-wild human motion video datasets and 3x larger than existing MOCAP ground-truth datasets, and is the first large-scale benchmark to provide synchronized multi-view recordings of expert athletes. The multi-view setup enables low-cost, label-free evaluation of monocular 3D pose estimation algorithms. The authors describe a standardized collection protocol requiring no specialized equipment, along with fully automated video calibration and synchronization. Benchmarking state-of-the-art methods reveals that while 3D joint angle recovery is now fairly accurate, all models struggle with consistent depth estimation and foot contact. Two new metrics — footwork and stability — are proposed, exposing previously unexplored failure modes and highlighting concrete improvement opportunities in pose estimation and motion analysis.

Paper Overview

Research area: Computer Vision Authors: Ilona Demler, Xinran Xie, Blake Werner Published: 2026-06-20 arXiv: 2506.16207

Abstract

The Caltech Tennis dataset (CalTennis) is a large-scale video benchmark for evaluating monocular-to-3D pose estimation in the wild. CalTennis contains over 11 million frames (51 hours) of tennis practice and matches from 40 players, captured with 2-6 synchronized cameras at 60Hz.

Key facts:

  • 10x larger than existing in-the-wild human motion video datasets, and 3x larger than existing MOCAP ground-truth datasets
  • The first large-scale benchmark to provide synchronized multi-view recordings of expert athletes
  • The multi-view setup enables low-cost, label-free evaluation of monocular-to-3D pose estimation algorithms
  • The authors present a simple, standardized data collection protocol requiring no specialized equipment or expertise, along with fully automated video calibration and synchronization

Findings

Benchmarking state-of-the-art monocular-to-3D pose estimation methods on CalTennis shows that while 3D joint angle recovery is now quite accurate, all models struggle to consistently estimate depth and foot contact. The authors further propose two new performance metrics — footwork and stability — along with a qualitative study of body-shape inconsistencies. These metrics reveal previously unexplored failure modes and point to concrete opportunities for improvement in pose estimation and motion analysis.

---

*Automatically collected on 2026-06-21*

Tags

#computer-vision#pose-estimation#dataset#benchmark#3d-pose#tennis#multi-view#motion-analysis

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981609