English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Sharp Asymptotics for Kernel Ridge Regression under Anisotropic Power-Law Data

Forum topic · 小凯 · 2026-09-01

Summary

This paper by Lorenzo Rizzi, Arie Wortsman Zurich, and Bruno Loureiro (arXiv:2608.28564) studies kernel ridge regression with anisotropic Gaussian data whose covariance decays as a power law with exponent α ≥ 0. Using polynomial inner-product kernels, the authors derive asymptotically sharp expressions for the kernel spectrum and generalization error in the polynomial high-dimensional regime n = Θ(d^κ), showing how anisotropy reshapes learning curves. For weak anisotropy (0 < α < 1), variance peaks at integer sample complexities are progressively damped as α grows, while bias transitions decouple from interpolation peaks at fractional sample complexities for well-aligned targets. For strong anisotropy (α > 1), the effective dimension becomes constant: variance becomes independent of sample size, and the bias undergoes sharp transitions—abrupt learning below a threshold governed by the target's decay rate, or classic power-law decay with source and capacity rates above it. Results are specialized to single-index targets, showing how alignment with principal directions determines anisotropy's effect on learning.

Paper Overview

  • Field: Machine Learning
  • Authors: Lorenzo Rizzi, Arie Wortsman Zurich, Bruno Loureiro
  • Published: 2026-08-28
  • arXiv: 2608.28564

Abstract

We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent \(α\geq 0\) for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime \(n=Θ(d^κ)\), revealing how anisotropy reshapes the learning curves.

For weak anisotropy (\(0<α<1\)), the problem remains effectively high-dimensional and retains some features of the isotropic case, while departing from it in others: the variance still peaks at integer sample complexities \(κ\in\mathbb{N}\), but these peaks are progressively damped as \(α\) grows; meanwhile, for targets strongly aligned with the data's principal directions, the bias drops at fractional sample complexities, decoupling the bias transition from the interpolation peaks.

For strong anisotropy (\(α>1\)), the problem has constant effective dimension: the variance no longer depends on sample size, plateauing under unregularized interpolation or vanishing at an explicit rate under a fixed ridge penalty. The bias undergoes a sharp transition controlled by the target's decay rate: below the threshold, learning is abrupt rather than gradual; above it, the bias decays as a power law, recovering the classical source and capacity rates.

Finally, the results are specialized to single-index targets, showing how the alignment between the index and the data's principal directions determines how anisotropy affects learning.

---

*Auto-collected on 2026-09-01*

Tags

#machine-learning#kernel-ridge-regression#kernel-methods#high-dimensional-statistics#learning-theory#anisotropic-data#generalization-error#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634337