Summary
Tabular foundation models such as TabPFN and TabICL produce full predictive distributions, yet prevailing regression benchmarks evaluate them almost exclusively via point-estimate metrics like RMSE and R2. These aggregate measures can obscure model performance in the tails of the distribution—a critical deficit for high-stakes decision making in domains such as finance and clinical research, where asymmetric risk profiles are the norm. Researchers Jonas Landsgesell and Pascal Knoll introduce ScoringBench, an open benchmark that computes a comprehensive suite of proper scoring rules—including CRPS, CRLS, Interval Score, Energy Score, weighted CRPS, and Brier Score—alongside standard point metrics, providing a richer picture of probabilistic forecast quality. The paper evaluates realTabPFNv2.5 fine-tuned with different scoring-rule objectives, as well as TabICL, on a regression benchmark suite. Results confirm that model rankings depend on the chosen scoring rule, indicating that no single pre-training objective is universally optimal. The paper is available on arXiv (2603.11115), released March 31, 2026.
Paper Overview
Field: AI
Authors: Jonas Landsgesell, Pascal Knoll
Released: 2026-03-31
arXiv: 2603.11115
Abstract
Tabular foundation models such as TabPFN and TabICL already produce full predictive distributions, yet prevailing regression benchmarks evaluate them almost exclusively via point-estimate metrics such as RMSE and R2. These aggregate measures often obscure model performance in the tails of the distribution — a critical deficit for high-stakes decision making in domains like finance and clinical research, where asymmetric risk profiles are the norm.
The authors introduce ScoringBench, an open benchmark that computes a comprehensive suite of proper scoring rules — CRPS, CRLS, Interval Score, Energy Score, weighted CRPS, and Brier Score — alongside standard point metrics, providing a richer picture of probabilistic forecast quality.
Evaluation
The paper evaluates realTabPFNv2.5 fine-tuned with different scoring-rule objectives, as well as TabICL, relative to the unadjusted realTabPFNv2.5 baseline, on a suite of regression benchmarks.
Key Findings
- Model rankings depend on the chosen scoring rule.
- No single pre-training objective is universally optimal across all probabilistic evaluation metrics.
Links
- arXiv: https://arxiv.org/abs/2603.11115
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177169494