English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

Forum topic · 小凯 · 2026-08-27

Summary

A new arXiv paper (2608.24814) reveals the phenomenon of ELR collapse in language model pretraining: the learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across training runs, loss trajectories collapse throughout training despite substantially different LRs and parameter norms. This holds across optimizers, architectures, datasets, and model scales, with mean collapse errors typically on the order of a few times 10^-3 — below measured seed-to-seed variation in a representative configuration. Systematic ablations identify normalization design and the timescale of LR-norm variation as key determinants of collapse precision. Controlled interventions show that weight decay and HyperBall shape loss dynamics mainly through the ELR schedules they induce. Replacing LR with ELR enables fitted-function scaling laws (FSL) to transfer across norm-control methods. The paper was published on August 25, 2026, by Zihan Liu, Ruiheng Zheng, Shaobo Zhang, Changxin Tian, Kunlong Chen, and colleagues.

Paper Overview

Field: Machine Learning Authors: Zihan Liu, Ruiheng Zheng, Shaobo Zhang, Changxin Tian, Kunlong Chen, et al. Published: 2026-08-25 arXiv: 2608.24814

Key Findings

We uncover ELR collapse in language model pretraining: the learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio — the effective learning rate (ELR).

  • When ELR is matched across runs, their loss trajectories collapse throughout training, despite substantially different LRs and parameter norms.
  • Across optimizers, architectures, datasets, and model scales, mean collapse errors are typically a few times 10^-3, below the seed-to-seed variation measured in a representative configuration.
  • Systematic ablations identify normalization design and the timescale of LR–norm variation as key determinants of collapse precision.
  • Controlled interventions further show that weight decay and HyperBall shape loss dynamics primarily through the ELR schedules they induce.
  • Replacing LR with ELR enables fitted-function scaling laws (FSL) to transfer across norm-control methods.

Significance

This result suggests that ELR, rather than the raw learning rate, is the fundamental quantity controlling pretraining loss dynamics — with practical implications for hyperparameter transfer, optimizer design, and scaling-law extrapolation.

---

*Auto-collected on 2026-08-27*

Tags

#machine-learning#language-models#pretraining#learning-rate#scaling-laws#optimization#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634097