Paper Overview
- Field: Machine Learning
- Authors: Lorenzo Rizzi, Arie Wortsman Zurich, Bruno Loureiro
- Published: 2026-08-28
- arXiv: 2608.28564
Abstract
We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent \(α\geq 0\) for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime \(n=Θ(d^κ)\), revealing how anisotropy reshapes the learning curves.
For weak anisotropy (\(0<α<1\)), the problem remains effectively high-dimensional and retains some features of the isotropic case, while departing from it in others: the variance still peaks at integer sample complexities \(κ\in\mathbb{N}\), but these peaks are progressively damped as \(α\) grows; meanwhile, for targets strongly aligned with the data's principal directions, the bias drops at fractional sample complexities, decoupling the bias transition from the interpolation peaks.
For strong anisotropy (\(α>1\)), the problem has constant effective dimension: the variance no longer depends on sample size, plateauing under unregularized interpolation or vanishing at an explicit rate under a fixed ridge penalty. The bias undergoes a sharp transition controlled by the target's decay rate: below the threshold, learning is abrupt rather than gradual; above it, the bias decays as a power law, recovering the classical source and capacity rates.
Finally, the results are specialized to single-index targets, showing how the alignment between the index and the data's principal directions determines how anisotropy affects learning.
---
*Auto-collected on 2026-09-01*