English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

Forum topic · 小凯 · 2026-09-10

Summary

This paper (arXiv:2609.09116) by Hasan Amin, Wei-Kai Chang, and Rajiv Khanna studies why scale-invariant neural networks—rendered so by normalization—become unstable under learning-rate schedules and weight decay. The authors derive an exact discrete-time law in which a single scalar quantity captures all schedule and decay forcing, while norm growth induces an opposing geometric self-quenching effect. This produces a sharp boundary separating contraction- and expansion-dominated effective learning-rate regimes. Through exact analysis of a fully solved normalized regression model whose dynamics reduce to two dimensions, they show the balance point is intrinsically unstable: constant learning rates with weight decay cannot maintain an interior equilibrium and instead yield recurrent behavior driven by discrete-time Jacobian structure. A unified homogeneous-optimizer framework reveals a structural dichotomy in self-quenching strength, explaining from first principles why adaptive optimizers stabilize more weakly under normalization. Experiments on MLPs, CNNs, and GPT2 across MNIST, CIFAR, wikiText, and OpenWebText confirm the law with high precision, with performance peaking sharply at the predicted boundary.

Paper Overview

  • Field: cs.LG
  • Authors: Hasan Amin, Wei-Kai Chang, Rajiv Khanna
  • Posted: 2026-09-08
  • arXiv: 2609.09116

Abstract

Normalization renders large parts of neural networks effectively scale invariant, inducing a hidden feedback loop in which learning-rate schedules and weight decay interact through the parameter norm to control the effective step taken by the optimizer. We show that this interaction is governed by an exact discrete-time law: a single scalar quantity captures all schedule and decay forcing, while norm growth induces an opposing geometric self-quenching effect. This yields a sharp boundary that cleanly separates contraction- and expansion-dominated effective learning rate regimes. To understand the underlying mechanism, we provide exact analysis of a fully solved normalized regression model where the dynamics reduce to two dimensions and show that the balance point is intrinsically unstable, implying that constant learning rate with weight decay cannot stably maintain an interior equilibrium and instead produces recurrent behavior driven by discrete-time Jacobian structure. We further extend this perspective across optimizers through unified homogeneous-optimizer framework that reveals a structural dichotomy in self-quenching strength, providing a first-principles explanation for why adaptive methods exhibit systematically weaker stabilization under normalization. Across dynamical systems and neural networks (MLP, CNN, GPT2 / MNIST, CIFAR, wikiText, OpenWebText), the predicted law holds with high precision and enables direct control of training via the identified scalar, with performance peaking sharply at the predicted boundary. Together, these results isolate a single governing quantity for scale-invariant optimization, providing a precise and actionable lens on training dynamics, optimizer behavior, and schedule design in modern deep learning.

--- *Auto-collected on 2026-09-10*

Tags

#deep-learning#optimization#scale-invariance#weight-decay#learning-rate-schedule#training-dynamics#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634682