English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Device Data

Forum topic · 小凯 · 2026-09-08

Summary

WearableQA is a new benchmark (arXiv:2609.05405) for evaluating whether AI systems can reason over real users' longitudinal wearable records. The dataset contains 4,084 ten-option multiple-choice questions built from wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. Unlike prior benchmarks, WearableQA preserves authentic wearable data distributions, including device noise and inter-individual variability. It defines 16 question types organized along two complementary axes: data versus health reasoning, distinguishing numerical computation over longitudinal measurements from physiological interpretation. The work is authored by Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, and Benoit Corda, and falls within the NLP field. Details and the abstract are available on arXiv.

Paper Overview

Field: NLP Authors: Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda arXiv: 2609.05405

Abstract (from arXiv)

Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation. (Abstract truncated in source post; see arXiv page for the full version.)

Key Facts

  • 4,084 multiple-choice questions with 10 options each
  • Data from 200 real users with up to 500 days of daily wearable measurements
  • Includes wearable time series, blood biomarkers, and demographics
  • 16 question types along two axes: data vs. health reasoning
  • Preserves realistic device noise and inter-individual variability
For the full paper, see: https://arxiv.org/abs/2609.05405

Tags

#wearableqa#benchmark#nlp#wearables#health-reasoning#arxiv#llm-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634618