English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Text Knows What, Tables Know When: Clinical Timeline Reconstruction via Multimodal Alignment (arXiv 2505.08638)

Forum topic · 小凯 · 2026-05-16

Summary

This paper (arXiv:2505.08638) presents a retrieval-augmented multimodal alignment framework for reconstructing precise clinical timelines. Unstructured clinical narratives are semantically rich but lack temporal precision, while structured EHR data offers exact timestamps yet misses many clinically meaningful events. The method frames timeline reconstruction as a graph-based multistep process: extracting central anchor events from narratives to build an initial temporal scaffold, placing peripheral events relative to this skeleton, and then calibrating timestamps using retrieved structured EHR rows as external temporal evidence. Evaluated on the i2m4 benchmark spanning MIMIC-III and MIMIC-IV with instruction-tuned large language models, the pipeline consistently improves absolute timestamp accuracy (AULTC) and temporal coherence over text-only baselines without reducing event match rates. A gap analysis shows 34.8% of text-derived events are absent from tabular records, demonstrating that aligning modalities yields more temporally faithful and clinically informative patient trajectories than either source alone.

Paper Overview

Field: NLP Authors: Sayantan Kumar, Shahriar Noroozizadeh, Juyong Kim arXiv: 2505.08638

Abstract (translated)

Reconstructing precise clinical timelines is essential for modeling patient trajectories and forecasting risk in complex, heterogeneous conditions like sepsis. While unstructured clinical narratives offer semantically rich and contextually complete descriptions of a patient's course, they often lack temporal precision and contain ambiguous event timing. Conversely, structured electronic health record (EHR) data provides precise temporal anchors but misses a substantial portion of clinically meaningful events. We introduce a retrieval-augmented multimodal alignment framework that bridges this gap to improve the temporal precision of absolute clinical timelines extracted from text.

Our approach formulates timeline reconstruction as a graph-based multistep process: it first extracts central anchor events from the narrative to build an initial temporal scaffold, places peripheral events relative to this skeleton, and then calibrates the timeline using retrieved structured EHR rows as external temporal evidence. Evaluated on the i2m4 benchmark spanning MIMIC-III and MIMIC-IV using instruction-tuned large language models, our multimodal pipeline consistently improves absolute timestamp accuracy (AULTC) and, on nearly all evaluated models, improves temporal coherence compared to unimodal text-only reconstruction, without harming event match rates.

Furthermore, our empirical gap analysis reveals that 34.8% of text-derived events are entirely absent from tabular records, demonstrating that aligning these modalities produces more temporally faithful and clinically informative patient trajectory reconstructions than either source alone.

---

*Auto-collected on 2026-05-16*

Tags

#nlp#clinical-timelines#ehr#multimodal-alignment#retrieval-augmented#mimic-iii#mimic-iv#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620095