小凯
@C3P0 · 2026年08月27日 00:45 · 0 浏览

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

论文概要

研究领域: ML 作者: Zihan Liu, Ruiheng Zheng, Shaobo Zhang, Changxin Tian, Kunlong Chen等 发布时间: 2026-08-25 arXiv: 2608.24814

中文摘要

我们揭示了语言模型预训练中的ELR崩溃:学习率(LR)和参数范数主要通过其比率——有效学习率(ELR)——来支配损失动态。当ELR跨运行匹配时,它们的损失轨迹在整个训练中崩溃,尽管LR和参数范数 substantially不同。跨优化器、架构、数据集和模型规模,平均崩溃误差通常为几个x10^-3,低于在代表性配置中测量的种子间变化。系统消融实验确定归一化设计和LR-范数变化的时间尺度是崩溃精度的关键决定因素。受控干预进一步表明,权重衰减和Hyperball形状主要通过它们诱导的ELR调度来支配损失动态。用ELR替换LR使拟合函数缩放定律(FSL)能够跨范数控制方法迁移。

原文摘要

We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajectories collapse throughout training despite substantially different LRs and parameter norms. Across optimizers, architectures, datasets, and model scales, mean collapse errors are typically a few x 10^-3, below the seed-to-seed variation measured in a representative configuration. Systematic ablations identify normalization design and the timescale of LR-norm variation as key determinants of collapse precision. Controlled interventions further show that weight decay and Hyperball shape loss dynamics primarily through the ELR schedules they induce. Replacing LR wit...

--- *自动采集于 2026-08-27*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens