English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

The SGD-Adam Gap in LLM Pretraining Is About Learning Rate, Not Adaptivity

Forum topic · 小凯 · 2026-05-19

Summary

Why does SGD dramatically underperform Adam when pretraining large language models? A paper by Glentis, Li, Yau, and Hong challenges the common belief that Adam's per-parameter adaptive learning rates are the key advantage. The authors argue the real cause is that LLM pretraining involves small gradient norms and large weight-to-gradient ratios, especially at large batch sizes, requiring a very large effective learning rate for meaningful weight updates. Adam naturally delivers this because its update magnitude is normalized by gradient scale, while SGD cannot. The bottleneck for SGD comes from the output layer, where rare tokens produce sudden gradient spikes that force a small maximum learning rate or risk training divergence. With simple gradient clipping, SGD closes the validation loss gap to Adam from over 50% to just 3.5% on a 1B-parameter LLaMA-style pretraining run, revealing that large effective learning rates—long enjoyed silently by Adam—explain most of the gap. Open questions include the source of the residual 3.5% gap, automatic selection of clipping thresholds, and whether findings hold at smaller batch sizes.

Why does SGD perform far worse than Adam when pretraining large language models? The widely accepted explanation is Adam's adaptive learning rate mechanism—adjusting the step size for each parameter—gives it an essential advantage in the high-dimensional, sparse-gradient setting of LLMs. Glentis, Li, Yau, and Hong revisit this assumption and find the answer may be simpler.

The real story: effective learning rate

LLM pretraining is characterized by small gradient norms and large weight-to-gradient ratios—an effect even more pronounced with large batch sizes. This means you need a very large effective learning rate for weight updates to change anything meaningfully. Adam's update magnitude is not limited by the gradient norm: its gradient-divided-by-gradient-norm mechanism naturally amplifies the effective step size. SGD has no such mechanism.

The obstacle: output-layer gradient spikes

The problem lies in the uneven gradient distribution at the output layer. Gradient magnitudes vary enormously across token classes—common words like "the" and "a" produce small gradients, while rare words can suddenly produce huge gradient spikes in certain batches. These spikes cap the maximum learning rate SGD can tolerate: push the learning rate slightly higher and a single spike blows up the entire model.

Fixing it with gradient clipping

With simple gradient clipping, when pretraining a 1B-parameter LLaMA model on a large dataset, SGD closes the validation loss gap to Adam from over 50% down to just 3.5%. Clipping lets SGD run stably at the large learning rates that Adam has always enjoyed without anyone noticing—that large effective learning rate, not adaptivity itself, explains most of the gap.

Open questions

  • What explains the remaining 3.5% gap—other benefits of adaptive learning rates, such as handling differences in parameter scale?
  • How should the clipping threshold be determined automatically?
  • The paper uses very large batches (1M tokens). Do these findings hold in more common smaller-batch settings?
---

References

1. Glentis, A., Li, D., Yau, C., & Hong, M. (2026). *Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates*. arXiv:2605.17787 [cs.LG]. 2. Kingma, D. P., & Ba, J. (2015). *Adam: A Method for Stochastic Optimization*. ICLR. 3. Zhang, J., et al. (2020). *Why Gradient Clipping Accelerates Training: A Theoretical Justification for Adaptivity*. ICLR.

Tags

#deep-learning#optimization#sgd#adam#llm-pretraining#gradient-clipping#learning-rate#research

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620382