English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

BERT-LER: Explainable Transformer Models for Clinical Prediction on Structured EHR Data

Forum topic · 小凯 · 2026-08-22

Summary

A forum post introduces BERT-LER, a BERT-style encoder for electronic health record (EHR) timelines presented in arXiv paper 2608.20315 by Jun Ni Du, Lukas Adamek, Maxim Kryukov, and Flavio Dormont. The model is pre-trained and fine-tuned on a de-identified EHR dataset covering 75 million patients. Its key design encodes laboratory test results as discrete tokens while preserving graded information through percentile-based binning, and it integrates Integrated Gradients to provide token-level attributions relative to the input medical event sequence. The model was evaluated on the public EHRShot benchmark and on an asthma severity progression study using real-world data. Results show competitive predictive performance, often outperforming publicly available baseline models on lab-related tasks, with attributions that align with clinically known risk factors. The architecture and interpretability approach are broadly applicable across therapeutic areas and prediction tasks.

Paper Overview

Field: Machine Learning Authors: Jun Ni Du, Lukas Adamek, Maxim Kryukov, Flavio Dormont Published: 2026-08-22 arXiv: 2608.20315

Abstract

Prediction models on structured electronic health records (EHR) remain central to medical machine learning, yet few approaches jointly emphasize quantitative laboratory information and explainability relative to input medical events. This paper proposes BERT-LER, a BERT-style model encoding EHR timelines, pre-trained and fine-tuned on a de-identified EHR dataset of 75 million patients.

The model encodes laboratory test results as discrete tokens while preserving graded information via percentile-based binning. Combined with Integrated Gradients, it provides token-level attributions grounded in the input EHR sequence.

Evaluation

  • EHRShot benchmark: competitive predictive performance against public baselines
  • Asthma severity progression study (real-world data): strong results on lab-related tasks
  • Explainability: token-level attributions align with clinically known risk factors

Conclusion

BERT-LER typically outperforms publicly available baseline models on lab-related tasks while offering clinically meaningful explanations. The architecture and interpretability methodology can be applied to many therapeutic areas and prediction tasks.

--- *Automatically collected on 2026-08-22*

Tags

#machine-learning#ehr#transformers#explainability#clinical-prediction#bert#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633817