Paper Overview
Research Area: Computer Vision (CV) Authors: Mario Tuci, Caner Korkmaz, Umut Şimşekli, Tolga Birdal Published: 2026-04-21 arXiv: 2604.19740
Summary
Training modern neural networks typically relies on large learning rates, operating at the edge of stability, where the optimization dynamics exhibit oscillatory and chaotic behavior. Empirically, this regime often yields improved generalization performance, yet the underlying mechanism remains poorly understood.
In this work, the authors represent stochastic optimizers as random dynamical systems, which often converge to a fractal attractor set (rather than a single point) with a smaller intrinsic dimension. Building on this connection and inspired by Lyapunov dimension theory, they introduce a novel notion of dimension, coined the sharpness dimension, and prove a generalization bound based on this dimension.
Key Contributions
- Random dynamical systems view: Stochastic optimizers are shown to converge to fractal attractor sets rather than single points, with smaller intrinsic dimension.
- Sharpness dimension: A new dimension concept inspired by Lyapunov dimension theory, used to derive generalization bounds.
- Complete Hessian spectrum: Generalization in the chaotic regime depends on the full Hessian spectrum and the structure of partial determinants—complexity that prior work based on trace or spectral norm fails to capture.
- Experimental validation: The theory is validated on a variety of MLPs and Transformers, and provides new insights into the recently observed grokking phenomenon.