Understanding Reasoning from Pretraining to Post-Training: What Chess Reveals About LLM Learning
Paper: Understanding Reasoning from Pretraining to Post-Training Authors: Jingyan Shen, Ang Li, Salman Rahman, et al. arXiv: 2607.16097 Published: 2026-07-17
Overview
Nearly a decade after AlphaZero defeated the world chess champion, a fundamental question still haunts AI researchers: how do large language models actually learn to *reason*? DeepSeek-R1 and OpenAI o1 can do it—but what did pretraining plant, and what did reinforcement learning (RL) grow?
This paper takes an elegant approach: instead of proving theorems in abstract math spaces, it returns to chess—a controlled, verifiable, fully understood domain—and builds a miniature LLM training pipeline to answer two questions:
1. How do pretraining choices (model size, data volume) determine the returns from RL? 2. Does RL sharpen existing capabilities, or discover entirely new ones?
The Three-Tier Pyramid: Standard LLM Training
1. Pretraining — self-supervised next-token prediction on massive corpora. For chess: learning statistical patterns from millions of human games (e.g., 1. e4 e5 2. Nf3 Nc6 3. Bb5 ...). Like a child absorbing language and common sense through extensive reading.
2. Supervised Fine-Tuning (SFT) — learning the surface structure of reasoning: positions paired with human-style rationales and moves. The model mimics *how* humans think, without understanding why.
3. Reinforcement Learning — playing actual games with win/loss rewards. No teacher labels moves; the model judges by outcomes. This is where AlphaZero-style discoveries can emerge.
Experimental Design
- Model sizes: 7 scales, from 5M to 1B parameters
- Pretraining data: 10M to 1B tokens
- Pipeline: pretraining on Human Chess Games → SFT on synthetic reasoning traces → RL with verifiable rewards on Chess Puzzles (unique checkable answers)
- Given a pretrained model, its pretraining loss nearly linearly predicts its final post-RL performance:
RL final performance ≈ a × pretraining loss + b. - More pretraining tokens → steeper RL learning curves: better-pretrained models are not only better but learn faster during RL.
- Metaphor from the post: pretraining is the river's source elevation; RL builds dams downstream—it can speed the flow but cannot raise the source.
- On easy puzzles (solvable after SFT), RL amplifies already-preferred correct moves—consistent with the "sharpening" view.
- On hard puzzles (SFT assigns near-zero probability to correct answers), RL systematically selects moves that SFT almost never produced. This is creation of new preferences, not amplification—echoing AlphaZero's startling openings that violated centuries of human chess theory.
- Better-pretrained checkpoints reach higher post-RL math ability
- RL curve slope scales roughly linearly with pretraining tokens
- RL discovers solution paths SFT never attempted on hard problems
- Pretraining is the foundation: RL cannot compensate for insufficient pretraining; quality control in pretraining matters more than RL tuning.
- Pretraining loss is an underrated metric: you can predict post-RL performance without running RL—valuable for checkpoint selection, early stopping, and compute allocation.
- RL's value is discovery, not just optimization: reward functions focused on final-answer correctness (rather than imitating human reasoning formats) may unlock non-human reasoning paths; over-constraining chain-of-thought format may limit RL's exploratory power.
- Cross-stage joint optimization: since pretraining loss encodes RL potential, future objectives could predictively maximize post-RL performance during pretraining itself.
- Fully deterministic rules, no ambiguity
- Huge (~10^44 positions) but verifiable state space
- Millennia of human knowledge for comparison
- Automatic verification of game outcomes and puzzle answers
Key findings
1. Pretraining loss predicts post-RL performance
2. RL discovers new capabilities, not just sharpening
3. Replication in mathematics
The same laws hold for a 1B-parameter math model (pretraining on math text → SFT with reasoning traces → RL with verifiable answers):
This suggests a cross-domain, domain-agnostic mapping between pretraining quality and RL outcomes.
Implications for LLM training
Why chess works as a testbed
References
1. Shen, J., Li, A., Rahman, S., et al. (2026). Understanding Reasoning from Pretraining to Post-Training. *arXiv:2607.16097* 2. Silver, D., et al. (2017). Mastering the Game of Go Without Human Knowledge. *Nature*, 550(7676), 354-359 3. Silver, D., et al. (2018). A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play. *Science*, 362(6419), 1140-1144 4. Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback. *NeurIPS 2022* 5. Guo, D., et al. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. *arXiv:2501.12948* 6. Kaplan, J., et al. (2020). Scaling Laws for Neural Language Models. *arXiv:2001.08361* 7. Hoffmann, J., et al. (2022). Training Compute-Optimal Large Language Models. *NeurIPS 2022* 8. Wei, J., et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. *NeurIPS 2022*