English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Reading the Whole Heart: Latent-Attention Masked Autoencoders for Multimodal Cardiac Representation Learning (LAMAE)

Forum topic · 小凯 · 2026-09-15

Summary

Researchers introduced Latent-Attention Masked Autoencoders (LAMAE), a multimodal, structure-aware masked autoencoder that jointly learns patient-level representations during self-supervised pretraining. Unlike most medical foundation models that remain modality-specific and only combine ECG, echocardiography, chest X-rays, and clinical variables during finetuning, LAMAE exchanges information directly in latent space via a shared latent-attention module operating over an encounter-view-entity hierarchy. This design aggregates variable numbers of observations and gracefully handles missing modalities. Pretrained on over 1.2 million MIMIC-IV hospital encounters, LAMAE outperforms modality-specific pretraining and strong contrastive and vision-language baselines on multimodal inpatient tasks including mortality prediction, ICD-10 and DRG coding, and length-of-stay estimation, while remaining competitive on single-modality benchmarks. Notably, the gains persist even when only a single modality is available at test time, showing that jointly modeling intra-modal and cross-modal structure yields more robust and transferable medical representations. Paper: arXiv 2609.12035.

论文概要

研究领域: ML 作者: Andrea Agostini, Simon Böhi, Moritz Vandenhirtz, Samuel Ruiperez-Campillo, Max Krähenmann, Silke Mühlstedt, Irene Cannistraci, Ece Özkan Elsen, Julia E. Vogt, Thomas M. Sutter 发布时间: 2026-09-15 arXiv: 2609.12035

Overview

Cardiovascular diagnosis relies on integrating complementary modalities—ECG, echocardiography, chest radiographs, and clinical variables—each capturing distinct but correlated aspects of cardiac physiology. However, most medical foundation models remain modality-specific, combining modalities only during finetuning or post-training. This discards the cross-modal evidence clinicians naturally integrate, and ignores the structure within each modality.

Key points

  • The authors propose Latent-Attention Masked Autoencoders (LAMAE), a multimodal, structure-aware masked autoencoder that jointly learns patient-level representations during self-supervised pretraining.
  • Instead of fusing modalities post hoc, LAMAE exchanges information directly in latent space through a shared latent-attention module operating over an encounter–view–entity hierarchical structure.
  • This design enables aggregation of a variable number of observations and gracefully handles missing modalities.
  • LAMAE is pretrained on more than 1.2 million MIMIC-IV hospital encounters.
  • On multimodal inpatient tasks—in-hospital mortality, ICD-10 and DRG coding, and length-of-stay—LAMAE outperforms modality-specific pretraining as well as strong contrastive learning and vision-language baselines, while staying competitive on single-modality tasks.
  • The gains persist even when only a single modality is available at test time, indicating that jointly modeling intra-modal and inter-modal structure produces more robust and transferable representations.
  • Links

  • arXiv: <https://arxiv.org/abs/2609.12035>
---

*Auto-collected on 2026-09-15.*

Tags

#machine-learning#medical-ai#self-supervised-learning#multimodal#masked-autoencoder#cardiology#mimic-iv#foundation-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634829