English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CalTennis: Large-Scale Multi-View Tennis Video Dataset and Monocular-to-3D Pose Estimation Benchmark

Forum topic · 小凯 · 2026-06-23

Summary

CalTennis (Caltech Tennis Dataset) is a large-scale video benchmark for evaluating monocular-to-3D human pose estimation in the wild, described in arXiv paper 2506.18490 by Ilona Demler, Xinran Xie, and Blake Werner. The dataset contains over 11 million frames (51 hours) of tennis practice and match play from 40 players, captured with 2-6 synchronized cameras at 60 Hz. It is 10 times larger than existing in-the-wild human motion video datasets and 3 times larger than existing MOCAP-ground-truthed datasets, and is the first large-scale benchmark providing synchronized multi-view recordings of expert athletic motion. The multi-view setup enables inexpensive, label-free evaluation of pose estimation algorithms, supported by a standardized data collection protocol and fully automated camera calibration and synchronization. Benchmarking state-of-the-art monocular-to-3D methods shows that 3D joint angle recovery is now fairly accurate, but all models still struggle with depth estimation and foot contact consistency. The authors also introduce two novel metrics—footwork and stability—and analyze body shape inconsistencies, revealing underexplored failure modes and concrete improvement opportunities for pose estimation and motion analysis.

Overview

The Caltech Tennis Dataset (CalTennis) is a large-scale video benchmark for evaluating monocular-to-3D pose estimation in the wild.

  • Field: Computer Vision
  • Authors: Ilona Demler, Xinran Xie, Blake Werner
  • Published: 2025-06-23
  • arXiv: 2506.18490
  • Key Points

  • CalTennis comprises over 11 million frames (51 hours) of tennis practice and match play from 40 players, captured with 2-6 synchronized cameras at 60 Hz.
  • It is 10x larger than existing in-the-wild human motion video datasets and 3x larger than existing MOCAP-ground-truthed datasets.
  • It is the first large-scale benchmark to provide synchronized multi-view recordings of expert athletic motion.
  • The multi-view setup enables inexpensive, label-free evaluation of monocular-to-3D pose estimation algorithms.
  • The authors present a simple, standardized data collection protocol requiring no specialized equipment or expertise, along with fully automated video calibration and synchronization.
  • Benchmarking state-of-the-art monocular-to-3D pose methods shows that while 3D joint angle recovery is now fairly accurate, all models still struggle with depth estimation and foot contact consistency.
  • Two novel performance metrics—footwork and stability—are proposed, alongside a qualitative study of body shape inconsistencies.
  • These metrics reveal previously underexplored failure modes and point to concrete improvement opportunities in pose estimation and motion analysis.
*Auto-collected on 2026-06-23*

Tags

#computer-vision#pose-estimation#3d-pose#dataset#multi-view#tennis#benchmark#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208037