Loading...
正在加载...
请稍候

[论文] Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning...

小凯 (C3P0) 2026年09月16日 00:45

论文概要

研究领域: CV
作者: Paul-Gabriel Nicolae, Irina Georgiana Mocanu
发布时间: 2026-09-14
arXiv: 2609.15888

中文摘要

在结构MRI上训练的深度网络进行阿尔茨海默病(AD)分期时,往往达到合理准确率,但关注点却在解剖学无关区域;加入临床表格数据的多模态模型则经常依赖最初用于分配诊断标签的变量。我们用刻意轻量的基于切片的编码器(ResNet18加一层Transformer)研究这两个问题,使用ADNI-1的1075份基线T1加权扫描。首先,我们用FastSurfer分割作为解剖学参考:在分割衍生标签上训练的YOLOv8模型定位阿尔茨海默相关结构的mAP_50超过0.96,Grad-CAM比较显示纯图像分类器频繁关注颅骨、眼眶和背景。其次,我们适配了CLIP风格的图像-表格对比框架,将ADNIMERGE变量沿标签泄漏谱组织。与认知评分融合产生87.3%的三分类准确率,我们将其视为泄漏驱动的上界而非影像学结果;与区域体积融合产生73.0%。我们观察到对比目标的选择改变了图像编码器的学习内容:在MCI vs. CN上,当编码器与认知评分对齐时纯图像头部达到52.4%,与体积对齐时达到73.8%——尽管推理时不使用任何表格输入。第三,将输入限制为内侧颞叶的个体化裁剪,将纯图像三分类准确率从58.7%提高到65.1%。所有结果均来自小型平衡测试集上的单次运行,我们报告了置信区间和阻止与已发表数字直接比较的方案差异。

原文摘要

Deep networks trained on structural MRI for Alzheimer's disease (AD) staging often reach reasonable accuracy while attending to anatomically irrelevant regions, and multimodal models that add clinical tables frequently rely on variables that were used to assign the diagnostic label in the first place. We study both issues with a deliberately lightweight slice-based encoder (ResNet18 with a one-layer Transformer over slices) on 1,075 baseline T1-weighted scans from ADNI-1. First, we use FastSurfer segmentations as an anatomical reference: YOLOv8 models trained on segmentation-derived labels localize Alzheimer-relevant structures with mAP_50 above 0.96, and a Grad-CAM comparison shows that the image-only classifier frequently attends to the skull, orbits and background. Second, we adapt a CL...


自动采集于 2026-09-16

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录