Skaling Laws: Chinchilla and Kaplan Were Both Half Right
> Paper: https://arxiv.org/abs/2608.07222 > Authors: Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz, Kartik Ahuja (FAIR at Meta)
A Four-Year-Old Debate
In 2020, OpenAI published Kaplan's Law: model loss decreases as a power law with parameter count N, data size D, and compute C. The key claim was that model size and data affect loss independently and can be optimized separately.
In 2022, DeepMind published Chinchilla's Law: Kaplan was wrong—parameters and data should scale proportionally, with an optimal token-to-parameter ratio of about 20:1. Chinchilla refit Kaplan's data and showed the independence assumption led Kaplan to severely underestimate the importance of data.
Since then, Chinchilla became the industry consensus. Open-source models like Llama, Mistral, and Qwen largely follow the 20:1 ratio.
But one problem has long been overlooked: Chinchilla's prediction error rises sharply in the extremes of data scarcity and over-training. Use Chinchilla to predict the loss of "10B parameters trained on 1T tokens" and it overestimates; predict "1B parameters trained on 1T tokens" and it underestimates.
The Skaling paper's opening sentence states the problem: the standard formula systematically under- and overestimates loss in the data-scarce and over-trained extremes. This is not random noise—it is systematic bias.
Root Cause: The Independence Assumption Is Wrong
Chinchilla's mathematical form is:
Here N and D appear as two independent terms, each decaying as a power law. That "independence" is the problem.
Skaling's core insight: the effects of model size and data volume on loss are not independent—they are coupled. The way a small model degrades on large data differs fundamentally from how a large model degrades on small data. A fixed \(A/N^\alpha\) term cannot describe all (N, D) combinations.
The authors demonstrate this with numerical gradient analysis: in the (N, D) plane, loss contours are not the "independent-sum" shape Chinchilla predicts, but clearly curved—the hallmark of coupling.
Skaling: Adding a Coupling Term
Skaling's form:
Compared to Chinchilla, it adds a cross term \(\frac{C}{(N^\alpha \cdot D^\beta)^\gamma}\) with coupling exponent \(\gamma\). The term means: as N and D grow together, there is a synergy (or offsetting effect) that independent terms cannot capture.
Why multiplicative coupling rather than additive? Section 7.2 of the paper argues in detail that additive interaction terms numerically fail to fit the observed contour curvature, while multiplicative coupling succeeds. This echoes Kaplan's original coupled form, but Skaling makes it more precise.
Results: 1.5-3x Error Reduction
Across multiple cross-validation settings, Skaling compared to Chinchilla:
- Interpolation (predicting points inside the training grid): MAPE reduced 1.5-2x
- Extrapolation (predicting outside the training grid): MAPE reduced 2-3x
- Sparse-grid strategy: a small number of training points from low-compute regions suffices to accurately extrapolate the whole grid—roughly 10x compute savings versus uniform sampling
- As \(\gamma \to 0\) (coupling vanishes), Skaling reduces to Chinchilla
- When data far exceeds what the model capacity needs, the coupling term dominates and Skaling behaves like Kaplan's coupled form
- Skaling adds a parameter \(\gamma\), making fitting harder than Chinchilla. With sparse data, the estimate of \(\gamma\) may itself be unstable.
- Experiments were done mainly on FAIR's internal training data. Whether \(\gamma\) is consistent across different data qualities and architectures needs more validation.
- The physical meaning of the coupling term is unclear. Why would N and D synergize? Because larger models extract more from data, or because more data better activates large models' capacity? The paper gives numerical evidence but no mechanistic explanation.
The third point matters most. LLM training costs run into millions of dollars; accurately predicting large-scale performance from small experiments saves enormous trial-and-error cost. Skaling's sparse-grid strategy makes "predicting a 7B model's loss from 100M-parameter experiments" feasible, with acceptable error.
The Deeper Insight: Both Kaplan and Chinchilla Were Half Right
The paper's most elegant point: Kaplan and Chinchilla are not opposed—they are approximations of Skaling in different limits.
Like relativity reducing to Newtonian mechanics at low speeds—neither is wrong, they have different domains of validity. Chinchilla is accurate in the "proportional scaling" sweet spot but fails at the extremes. Skaling unifies both extremes with one extra parameter.
Practical Implications for Training
1. Over-training finally has theoretical grounding. Llama 3 trains an 8B model on 15T tokens—far beyond Chinchilla's 20:1 ratio (which would imply 160B tokens). Such over-training looks "suboptimal" under Chinchilla, but Skaling explains it correctly: the coupling term makes the marginal loss of over-training smaller than Chinchilla predicts.
2. Small-model experiments predict large models more reliably. The sparse-grid strategy lets teams fit reliable scaling laws at 1/10 the compute—especially valuable for resource-constrained labs.
3. A new data-allocation frontier. The paper proposes a data allocation frontier: given total compute, how to split between N and D to minimize loss. Skaling's optimal allocation differs from Chinchilla's, especially at non-standard ratios.
An Honest Assessment
Limitations are clear:
---
Paper: Videau, Youbi-Idrissi, Lopez-Paz, Ahuja. *Skaling: Chinchilla's Exponents Meet Kaplan's Coupling*. arXiv:2608.07222, 2026. (FAIR at Meta)