English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Early-Stopped Negative-Shifted Gradient Descent

Forum topic · 小凯 · 2026-07-28

Summary

This forum post introduces arXiv paper 2607.22474 by Peng Zhao, which studies overparameterized linear regression. Many weak spectral directions act like an implicit ridge penalty on signal-bearing directions; negative ridge is a natural correction pushing filters above one, but stable negative-ridge endpoints are structurally limited: their pole must remain below the smallest nonzero empirical eigenvalue, and they anti-shrink small eigenvalues more than large ones. The paper shows early-stopped negative-shifted gradient descent escapes this constraint, producing a smooth filter at the would-be pole with mixed-sign capability: a leading prefix of above-ridgeless directions, with lower directions shrunk or exposure-controlled and stopping setting the crossover. In a Gaussian spike-plus-flat model, a Marchenko-Pastur barrier emerges: the shift canceling the implicit penalty lies one body-width above the smallest empirical eigenvalue, and the stopping path improves risk over every admissible endpoint by a polynomial factor. Main theorems allow general effective-rank tails, recovering all head scales and beating both positive shrinkage and uniform rescaling once scales separate. Localized Duhamel integrals handle the non-shrinking shift dynamics.

Paper Overview

  • Field: Machine Learning
  • Author: Peng Zhao
  • Published: 2026-07-24
  • arXiv: 2607.22474
  • Key Points

  • In overparameterized linear regression, many weak spectral directions act like a ridge penalty on the signal-bearing spectrum; the negative ridge is the natural correction, pushing filters above one.
  • The stable negative-ridge endpoint is structurally limited: its pole must stay below the smallest nonzero empirical eigenvalue, and it anti-shrinks smaller eigenvalues more than larger ones.
  • Early-stopped negative-shifted gradient descent escapes this constraint. Its filter is smooth at the would-be pole and mixed-sign-capable: above-ridgeless directions form a leading prefix, while lower directions are shrunk or exposure-controlled, with the stopping time setting the crossover.
  • In a Gaussian spike-plus-flat model, the paper discovers a Marchenko-Pastur barrier: the shift that cancels the implicit penalty lies one body-width above the smallest empirical eigenvalue, and the stopping path improves risk over every admissible endpoint by a polynomial factor under explicit conditions.
  • The main theorem allows general high effective-rank tails: the trace sets an implicit lower bound, the squared spectrum controls exposure, and the lower-bound critical path recovers all head scales simultaneously, surpassing positive shrinkage and, once scales separate, every uniform rescaling of the ridgeless solution.
  • The central technical challenge is handling the non-shrinking shift dynamics; localized Duhamel integrals are used to control them. A finite-grid retention inequality transfers the separations to validating algorithmic stopping choices.

Original Abstract (excerpt)

> In overparameterized linear regression, many weak spectral directions act like a ridge penalty on the signal-bearing spectrum; negative ridge is the natural correction, pushing filters above one. The stable negative-ridge endpoint, however, is structurally limited: its pole must stay below the smallest nonzero empirical eigenvalue, and it anti-shrinks smaller eigenvalues more than larger ones. Early-stopped negative-shifted gradient descent escapes this constraint. Its filter is smooth at the would-be pole and mixed-sign-capable: above-ridgeless directions form a leading prefix, with lower directions shrunk or exposure-controlled while stopping sets the crossover. In a Gaussian spike-plus-flat model we discover a Marchenko-Pastur barrier: the shift that cancels the implicit penalty lies a ...

---

*Auto-collected on 2026-07-28*

Tags

#machine-learning#linear-regression#gradient-descent#spectral-regularization#early-stopping#marchenko-pastur#arxiv#theory

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503750