Paper Overview
Field: NLP Authors: Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda arXiv: 2609.05405
Abstract (from arXiv)
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation. (Abstract truncated in source post; see arXiv page for the full version.)
Key Facts
- 4,084 multiple-choice questions with 10 options each
- Data from 200 real users with up to 500 days of daily wearable measurements
- Includes wearable time series, blood biomarkers, and demographics
- 16 question types along two axes: data vs. health reasoning
- Preserves realistic device noise and inter-individual variability