论文概要
研究领域: ML
作者: Andrea Agostini, Simon Böhi, Moritz Vandenhirtz, Samuel Ruiperez-Campillo, Max Krähenmann, Silke Mühlstedt, Irene Cannistraci, Ece Özkan Elsen, Julia E. Vogt, Thomas M. Sutter
发布时间: 2026-09-15
arXiv: 2609.12035
中文摘要
心血管诊断依赖于整合互补模态——心电图、超声心动图、胸部 X 光片与临床变量,每种模态捕捉心脏生理中不同但相关的侧面。然而大多数医学基础模型仍是特定模态的,仅在微调或训练后阶段才组合多模态。这丢弃了临床医生自然整合的跨模态证据,也忽略了每种模态内部的结构。我们提出潜在注意力掩码自编码器(LAMAE),一种多模态、结构感知的掩码自编码器,在自监督预训练期间联合学习患者级表征。LAMAE 不是在事后融合模态,而是通过一个作用于『检查-视图-实体』层级结构的共享潜在注意力模块,在潜在空间中直接交换信息,从而实现可变观测的聚合并优雅处理模态缺失。在超过 120 万次 MIMIC-IV 住院记录上预训练后,LAMAE 在住院死亡率、ICD-10 与 DRG 编码、住院时长等多模态住院任务上超越了特定模态预训练以及强对比学习与视觉-语言基线,同时在单模态任务上保持竞争力。即使测试时仅有单一模态可用,这些增益依然存在,表明同时对模态内与模态间结构建模能够产生更鲁棒、更可迁移的表征。
原文摘要
Cardiovascular diagnosis rests on integrating complementary modalities, like ECG, echocardiography, chest radiographs, and clinical variables, each capturing distinct but correlated aspects of cardiac physiology. Yet most medical foundation models remain modality-specific, combining modalities only for finetuning or post-training. This discards the cross-modal evidence clinicians naturally integrate and ignores the structure within each modality. We introduce Latent-Attention Masked Autoencoders (LAMAE), a multimodal, structure-aware masked autoencoder that jointly learns patient-level representations during self-supervised pretraining. Rather than fusing modalities post hoc, LAMAE exchanges information directly in the latent space through a shared latent-attention module operating over a ...
自动采集于 2026-09-15
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。