[论文] Transcriptome-informed multi-modal AI for predicting neoadjuvant thera...
研究领域: ML 作者: Jungkyu Park, Dhruva Biswas, Joseph Cappadona, Cerise Tang, Ken G. Zeng, Bartosz Machura, Chuwen Liu, Paolo Tarantino, Coral Omene, Francisco J. E…
论文概要
研究领域: ML 作者: Jungkyu Park, Dhruva Biswas, Joseph Cappadona, Cerise Tang, Ken G. Zeng, Bartosz Machura, Chuwen Liu, Paolo Tarantino, Coral Omene, Francisco J. Esteva, Rohit Bhargava, Marcin Braun, Kamila Paździerz, Jakub Czerwiński, Hanna Romańska-Knight, Albert Grinshpun, Bareket Daniel, Michele Buchinger, Frederick Howard, Piotr Wysocki, Brie Chun, Freya Schnabel, Rich Caruana, Jan Witowski, Krzysztof J. Geras 发布时间: 2026-10-02 arXiv: 2610.03693
中文摘要
标注数据的稀缺限制了肿瘤学深度学习生物标志物的发展。我们开发了一个两阶段 AI 模型,用于预测乳腺癌新辅助治疗的病理完全缓解(pCR)。第一阶段从组织病理学中学习转录组,使用 32 种癌症类型的 8742 名患者数据,经病理学家审查和与实测表达的空间一致性验证。这简化了第二阶段——从推断的表达和临床变量预测 pCR。模型使用 1080 名患者(五个队列)开发,在 1412 名患者(九个队列)中评估,合并 AUROC 为 0.79(95% CI,0.73-0.85),能在分子亚型内区分应答者。它优于组织病理学生物标志物,在瘤内采样和少量活检组织下保持稳定。消融实验表明,全转录组推断比仅用临床变量或单阶段病理模型具有更好的区分能力,并通过避免基因组检测的基因选择限制增强了鲁棒性。这些结果表明,生物学知情压缩可能推广到精准肿瘤学中数据稀疏的应用场景。
原文摘要
Scarcity of labeled data limits development of deep learning biomarkers in oncology. We develop a two-stage AI model predicting pathological complete response (pCR) to neoadjuvant therapy in breast cancer. The first stage learns the transcriptome from histopathology using 8,742 patients across 32 cancer types, corroborated by pathologist review and spatial agreement with measured expression. This simplifies the second stage to predicting pCR from inferred expression and clinical variables. Developed using 1,080 patients (five cohorts) and evaluated in 1,412 patients (nine cohorts), the model achieves a pooled AUROC of 0.79 (95% CI, 0.73-0.85), discriminating responders within molecular subtypes. It outperforms histopathological biomarkers, remaining stable across intratumoral sampling and ...
*自动采集于 2026-10-06*
#论文 #arXiv #ML #小凯