[论文] How abundant are good interpolators?
论文概要
研究领域: ML 作者: August Y. Chen, Ahmed El Alaoui 发布时间: 2026-06-04 arXiv: 2606.06469
中文摘要
设 S 为单位范数线性分类器 θ∈ℝ^d 的集合,这些分类器能够正确分类带标签数据集 (Xi,yi)_{i=1}^n 中的每个点,其中 Xi∈ℝ^d,yi∈{-1,+1},且预先固定了一个可能为负的间隔 κ。在 (X,y) 对的两种自然数据生成分布——高斯混合模型和具有高斯特征的逻辑模型——下,并在比例制度 n/d→α(α 足够小)中,我们建立了一个大偏差原理,用于描述在数据选择的高概率下,从 S 中均匀随机选取的点 θ 达到给定泛化误差的事件。相关的大偏差速率函数是确定性的,它描述了在 d 的指数尺度上,具有给定期望性能的插值分类器的比例。由此,我们建立了以下集中现象:除指数级小比例的插值分类器外,所有插值分类器都具有近似相同的泛化性能,该性能由该速率函数的唯一最大化值给出。我们将该最大化值与梯度下降的经验风险最小化性能以及一个自然线性规划的性能进行数值比较——两者都在 S 中寻找一个点——并推断,在小 α 的过参数化制度下,这些高效程序优于绝大多数插值器,表明在此设定中存在非平凡的良性过拟合。
原文摘要
Let S be the set of unit norm linear classifiers θ∈ ℝ^d which correctly classify every point of a labeled dataset (Xi,yi)_{i=1}^n, Xi ∈ ℝ^d, yi ∈ {-1,+1}, with a possibly negative margin κ fixed in advance. Under two natural data-generating distributions of the (X,y) pairs -- a Gaussian mixture model and a logistic model with Gaussian features -- and in the proportional regime n/d → α with small enough α, we establish a large deviation principle on the event that a point θ chosen uniformly at random from S achieves a given generalization error, with high probability over the choice of the data. The associated large deviation rate function is deterministic and describes the proportion, at the exponential scale in d, of interpolating classifiers having a given desired performance. As a conse...
--- *自动采集于 2026-06-08*
#论文 #arXiv #ML #小凯