English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Shannon Scaling Law: Modeling LLMs as Noisy Channels to Explain Overtraining and Quantization Degradation

Forum topic · 小凯 · 2026-05-26

Summary

A new paper (arXiv:2505.21433) by Xu Ouyang, Deyi Liu, and Yuhang Cai proposes the Shannon Scaling Law, a unified theoretical framework that models LLM training as information transmission over a noisy channel, grounded in the Shannon-Hartley theorem. By mapping model parameters to channel bandwidth and training tokens to signal power, the formulation captures the interplay between learning signal and intrinsic noise, revealing a fundamental Shannon capacity: scaling model size or data without preserving sufficient signal-to-noise ratio (SNR) amplifies noise, causing a transition from monotonic improvement to U-shaped performance degradation. This explains non-monotonic phenomena that classic power-law scaling laws cannot, such as catastrophic overtraining and quantization-induced degradation. Experiments on Pythia and OLMo2, spanning Gaussian noise, quantization, and supervised fine-tuning on math, QA, and code tasks, show the Shannon Scaling Law consistently outperforms classical and perturbation-aware laws with strong R² scores, accurately capturing loss basins. It also extrapolates: fitted on Pythia models up to 6.9B parameters and 180B tokens, it predicts unseen 12B models trained to 307B tokens with pooled R² = 0.847, where monotonic baselines fail entirely.

Paper Overview

Field: Machine Learning Authors: Xu Ouyang, Deyi Liu, Yuhang Cai Published: 2026-05-26 arXiv: 2505.21433

Key Idea

Existing scaling laws for LLMs are predominantly monotonic power laws and fail to explain emerging non-monotonic phenomena such as catastrophic overtraining and quantization-induced degradation, where performance deteriorates despite increased compute.

The authors propose the Shannon Scaling Law, a unified theoretical framework that models LLM training as information transmission over a noisy channel, grounded in the Shannon-Hartley theorem:

  • Model parameters are mapped to channel bandwidth
  • Training tokens are mapped to signal power
  • The formulation explicitly captures the interaction between learning signal and intrinsic noise
  • Implications

    This perspective reveals a fundamental Shannon capacity for LLMs: scaling model size or data without preserving a sufficient signal-to-noise ratio (SNR) inevitably amplifies noise, triggering a transition from monotonic improvement to U-shaped performance degradation.

    Experimental Validation

  • Evaluated on Pythia and OLMo2 models
  • Perturbations tested: Gaussian noise, quantization, and supervised fine-tuning on math, QA, and code tasks
  • The Shannon Scaling Law consistently outperforms classical scaling laws and recent perturbation-aware laws, achieving strong R² scores and accurately capturing loss basins that previous methods miss

Extrapolation Results

When fitted on Pythia models up to 6.9B parameters and 180B tokens, the law predicts unseen 12B models trained to 307B tokens with a pooled R² = 0.847, while monotonic baselines collapse entirely.

---

*Auto-collected on 2026-05-26*

Tags

#llm#scaling-laws#information-theory#shannon-capacity#overtraining#quantization#pythia#olmo2

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620811