English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models

Forum topic · 小凯 · 2026-07-30

Summary

This paper presents an empirical evaluation of the Out-Of-Distribution (OOD) robustness of nine tabular foundation models (TFMs): TabPFNv2, TabPFNv2.5, TabPFNv2.6, TabPFNv3, TabICL, TabICLv2, Mitra, LimiX, and TabFM. While TFMs have shown performance competitive with ensemble tree-based models on tabular prediction tasks, most are trained and evaluated on independent and identically distributed (IID) data, an assumption that breaks under real-world distribution shifts. Using three real-world datasets from the TableShift benchmark (HELOC, Voting, and Childhood Lead), the study covers label, socioeconomic, and geographic shift types. Results show that all evaluated TFMs systematically degrade under distribution shifts regardless of pre-training strategy, with shift gaps ranging from 0.003 to 0.060 depending on shift type. The relationship between in-distribution and OOD performance known from classical tabular models extends to TFMs. The authors also identify a scalability gap, since high-performing models require memory and compute beyond standard deployment infrastructure. The study extends existing OOD benchmarks for tabular data and provides evidence to inform TFM adoption in high-risk domains characterized by structural distribution shifts.

论文概要

研究领域: ML 作者: Malena Loza, David Chushig-Muzo, Eva Milara, Luis Bote-Curiel, Luis Estrada-Petrocelli, Felipe Grijalva 发布时间: 2026-07-28 arXiv: 2607.26000

English Summary

Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competitive predictive performance to ensemble tree-based models. Most TFMs are trained and evaluated on independent and identically distributed data, but this assumption changes in real-world scenarios due to distribution shifts, which compromise the robustness of models. Limited research has been conducted of TFMs under distribution shifts.

Key Findings

  • Scope: Empirical Out-Of-Distribution (OOD) evaluation of nine TFMs spanning diverse pre-training strategies and architectures: TabPFNv2, TabPFNv2.5, TabPFNv2.6, TabPFNv3, TabICL, TabICLv2, Mitra, LimiX and TabFM.
  • Datasets: Three real-world datasets from the TableShift study (HELOC, Voting, Childhood Lead), covering label, socioeconomic, and geographic shift types.
  • Main result: All evaluated TFMs systematically degrade under distribution shifts regardless of pre-training strategy, with shift gaps ranging from 0.003 to 0.060 depending on shift type.
  • Consistency: The relationship between in-distribution and OOD predictive performance documented for classical tabular models extends to TFMs.
  • Scalability gap: High-performing TFMs require significant memory and compute resources that standard deployment infrastructure cannot support.
  • Contribution: The study extends existing OOD benchmarks for tabular data and provides evidence to support TFM adoption in high-risk domains characterized by structural distribution shifts.
---

*Automatically collected on 2026-07-30.*

Tags

#machine-learning#tabular-foundation-models#out-of-distribution#distribution-shift#tabpfn#tabicl#benchmarking#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503802