Paper Overview
Field: Machine Learning Authors: Xu Ouyang, Deyi Liu, Yuhang Cai Published: 2026-05-26 arXiv: 2505.21433
Key Idea
Existing scaling laws for LLMs are predominantly monotonic power laws and fail to explain emerging non-monotonic phenomena such as catastrophic overtraining and quantization-induced degradation, where performance deteriorates despite increased compute.
The authors propose the Shannon Scaling Law, a unified theoretical framework that models LLM training as information transmission over a noisy channel, grounded in the Shannon-Hartley theorem:
- Model parameters are mapped to channel bandwidth
- Training tokens are mapped to signal power
- The formulation explicitly captures the interaction between learning signal and intrinsic noise
- Evaluated on Pythia and OLMo2 models
- Perturbations tested: Gaussian noise, quantization, and supervised fine-tuning on math, QA, and code tasks
- The Shannon Scaling Law consistently outperforms classical scaling laws and recent perturbation-aware laws, achieving strong R² scores and accurately capturing loss basins that previous methods miss
Implications
This perspective reveals a fundamental Shannon capacity for LLMs: scaling model size or data without preserving a sufficient signal-to-noise ratio (SNR) inevitably amplifies noise, triggering a transition from monotonic improvement to U-shaped performance degradation.
Experimental Validation
Extrapolation Results
When fitted on Pythia models up to 6.9B parameters and 180B tokens, the law predicts unseen 12B models trained to 307B tokens with a pooled R² = 0.847, while monotonic baselines collapse entirely.
---
*Auto-collected on 2026-05-26*