English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Generalization at the Edge of Stability: Sharpness Dimension and Fractal Attractors

Forum topic · 小凯 · 2026-04-23

Summary

This paper, 'Generalization at the Edge of Stability' by Mario Tuci, Caner Korkmaz, Umut Şimşekli, and Tolga Birdal (arXiv:2604.19740), explains why training neural networks with large learning rates at the edge of stability often improves generalization. The authors model stochastic optimizers as random dynamical systems that converge to fractal attractor sets rather than single points, with a smaller intrinsic dimension. Inspired by Lyapunov dimension theory, they introduce a new quantity called the 'sharpness dimension' and prove a generalization bound based on it. Their results show that generalization in the chaotic regime depends on the complete Hessian spectrum and the structure of partial determinants—complexity that trace- or spectral-norm-based analyses cannot capture. Experiments across MLPs and Transformers validate the theory and offer new insights into the grokking phenomenon.

Paper Overview

Research Area: Computer Vision (CV) Authors: Mario Tuci, Caner Korkmaz, Umut Şimşekli, Tolga Birdal Published: 2026-04-21 arXiv: 2604.19740

Summary

Training modern neural networks typically relies on large learning rates, operating at the edge of stability, where the optimization dynamics exhibit oscillatory and chaotic behavior. Empirically, this regime often yields improved generalization performance, yet the underlying mechanism remains poorly understood.

In this work, the authors represent stochastic optimizers as random dynamical systems, which often converge to a fractal attractor set (rather than a single point) with a smaller intrinsic dimension. Building on this connection and inspired by Lyapunov dimension theory, they introduce a novel notion of dimension, coined the sharpness dimension, and prove a generalization bound based on this dimension.

Key Contributions

  • Random dynamical systems view: Stochastic optimizers are shown to converge to fractal attractor sets rather than single points, with smaller intrinsic dimension.
  • Sharpness dimension: A new dimension concept inspired by Lyapunov dimension theory, used to derive generalization bounds.
  • Complete Hessian spectrum: Generalization in the chaotic regime depends on the full Hessian spectrum and the structure of partial determinants—complexity that prior work based on trace or spectral norm fails to capture.
  • Experimental validation: The theory is validated on a variety of MLPs and Transformers, and provides new insights into the recently observed grokking phenomenon.
--- *Auto-collected on 2026-04-23*

Tags

#deep-learning#optimization#edge-of-stability#generalization#chaotic-dynamics#sharpness-dimension#grokking#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618647