[论文] Reading the Whole Heart: Latent-Attention Masked Autoencoders for Mult...

研究领域: ML 作者: Andrea Agostini, Simon Böhi, Moritz Vandenhirtz, Samuel Ruiperez-Campillo, Max Krähenmann, Silke Mühlstedt, Irene Cannistraci, Ece Özkan Elsen, Ju…

论文概要

研究领域: ML 作者: Andrea Agostini, Simon Böhi, Moritz Vandenhirtz, Samuel Ruiperez-Campillo, Max Krähenmann, Silke Mühlstedt, Irene Cannistraci, Ece Özkan Elsen, Julia E. Vogt, Thomas M. Sutter 发布时间: 2026-09-15 arXiv: 2609.12035

中文摘要

心血管诊断依赖于整合互补模态——心电图、超声心动图、胸部 X 光片与临床变量,每种模态捕捉心脏生理中不同但相关的侧面。然而大多数医学基础模型仍是特定模态的,仅在微调或训练后阶段才组合多模态。这丢弃了临床医生自然整合的跨模态证据,也忽略了每种模态内部的结构。我们提出潜在注意力掩码自编码器(LAMAE),一种多模态、结构感知的掩码自编码器,在自监督预训练期间联合学习患者级表征。LAMAE 不是在事后融合模态,而是通过一个作用于『检查-视图-实体』层级结构的共享潜在注意力模块,在潜在空间中直接交换信息,从而实现可变观测的聚合并优雅处理模态缺失。在超过 120 万次 MIMIC-IV 住院记录上预训练后,LAMAE 在住院死亡率、ICD-10 与 DRG 编码、住院时长等多模态住院任务上超越了特定模态预训练以及强对比学习与视觉-语言基线,同时在单模态任务上保持竞争力。即使测试时仅有单一模态可用,这些增益依然存在,表明同时对模态内与模态间结构建模能够产生更鲁棒、更可迁移的表征。

原文摘要

Cardiovascular diagnosis rests on integrating complementary modalities, like ECG, echocardiography, chest radiographs, and clinical variables, each capturing distinct but correlated aspects of cardiac physiology. Yet most medical foundation models remain modality-specific, combining modalities only for finetuning or post-training. This discards the cross-modal evidence clinicians naturally integrate and ignores the structure within each modality. We introduce Latent-Attention Masked Autoencoders (LAMAE), a multimodal, structure-aware masked autoencoder that jointly learns patient-level representations during self-supervised pretraining. Rather than fusing modalities post hoc, LAMAE exchanges information directly in the latent space through a shared latent-attention module operating over a ...


*自动采集于 2026-09-15*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens