Paper Overview
- Field: Computer Vision / Medical Imaging
- Authors: Qi Chen, Wenxuan Li, Pedro R. A. S. Bassi
- arXiv: 2506.14717
- Current state-of-the-art AI models are optimized for average accuracy but perform worse on rare or underrepresented subgroups, such as young African American women.
- Performance also degrades on small tumors and non-standard imaging protocols.
- Collecting enough annotated data for these rare cases is often impractical, motivating systematic subgroup-level benchmarking instead.
Abstract
Artificial intelligence (AI) has achieved remarkable success in medical imaging, but it is widely recognized that these models often perform inconsistently across real-world clinical settings. Such inconsistencies occur when patient demographics and imaging protocols vary, for example, in detecting small tumors, analyzing scans from different contrast phases, or evaluating patients of different ages or sexes.
To quantify these inconsistencies, the authors develop a large-scale, open benchmark of 85,355 CT scans that systematically evaluates 12 tumor-detection AI models across tumor size, location, patient subgroup, and imaging protocol. They leverage large language models (LLMs) to extract and organize subgroup information from clinical data, which makes the analysis both scalable and reproducible.
Key Findings
Significance
The benchmark provides a foundation for building more reliable and robust tumor-detection AI models and highlights the need for rigorous subgroup-level evaluation in medical imaging and computer vision.
---
*Source: zhichai.net, auto-collected 2026-06-25.*