English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Understanding Reasoning from Pretraining to Post-Training: What Chess Reveals About LLM Learning

Forum topic · 小凯 · 2026-07-20

Summary

A forum post discusses the paper 'Understanding Reasoning from Pretraining to Post-Training' (arXiv:2607.16097), which uses chess as a controlled, verifiable testbed to study how large language models acquire reasoning. Training models from 5M to 1B parameters through a full pipeline of pretraining, supervised fine-tuning (SFT), and reinforcement learning (RL) with verifiable rewards, the authors find two key results. First, pretraining loss nearly linearly predicts final post-RL performance, and more pretraining tokens yield steeper RL learning curves—meaning pretraining quality sets an upper bound on RL gains before RL is ever run. Second, RL does not merely sharpen SFT policies: on hard puzzles where SFT assigns near-zero probability to correct moves, RL systematically discovers new solution strategies, echoing AlphaZero's superhuman openings. The same patterns replicate in a 1B-parameter math model, suggesting a domain-general law. Practical implications include using pretraining loss as a cheap potential indicator for checkpoint selection and resource allocation, and designing RL reward functions that reward correctness rather than human-imitative reasoning formats. The post ends by advocating controlled testbeds for studying AI fundamentals.

Understanding Reasoning from Pretraining to Post-Training: What Chess Reveals About LLM Learning

Paper: Understanding Reasoning from Pretraining to Post-Training Authors: Jingyan Shen, Ang Li, Salman Rahman, et al. arXiv: 2607.16097 Published: 2026-07-17

Overview

Nearly a decade after AlphaZero defeated the world chess champion, a fundamental question still haunts AI researchers: how do large language models actually learn to *reason*? DeepSeek-R1 and OpenAI o1 can do it—but what did pretraining plant, and what did reinforcement learning (RL) grow?

This paper takes an elegant approach: instead of proving theorems in abstract math spaces, it returns to chess—a controlled, verifiable, fully understood domain—and builds a miniature LLM training pipeline to answer two questions:

1. How do pretraining choices (model size, data volume) determine the returns from RL? 2. Does RL sharpen existing capabilities, or discover entirely new ones?

The Three-Tier Pyramid: Standard LLM Training

1. Pretraining — self-supervised next-token prediction on massive corpora. For chess: learning statistical patterns from millions of human games (e.g., 1. e4 e5 2. Nf3 Nc6 3. Bb5 ...). Like a child absorbing language and common sense through extensive reading. 2. Supervised Fine-Tuning (SFT) — learning the surface structure of reasoning: positions paired with human-style rationales and moves. The model mimics *how* humans think, without understanding why. 3. Reinforcement Learning — playing actual games with win/loss rewards. No teacher labels moves; the model judges by outcomes. This is where AlphaZero-style discoveries can emerge.

Experimental Design

  • Model sizes: 7 scales, from 5M to 1B parameters
  • Pretraining data: 10M to 1B tokens
  • Pipeline: pretraining on Human Chess Games → SFT on synthetic reasoning traces → RL with verifiable rewards on Chess Puzzles (unique checkable answers)
  • Key findings

    1. Pretraining loss predicts post-RL performance

  • Given a pretrained model, its pretraining loss nearly linearly predicts its final post-RL performance: RL final performance ≈ a × pretraining loss + b.
  • More pretraining tokens → steeper RL learning curves: better-pretrained models are not only better but learn faster during RL.
  • Metaphor from the post: pretraining is the river's source elevation; RL builds dams downstream—it can speed the flow but cannot raise the source.
  • 2. RL discovers new capabilities, not just sharpening

  • On easy puzzles (solvable after SFT), RL amplifies already-preferred correct moves—consistent with the "sharpening" view.
  • On hard puzzles (SFT assigns near-zero probability to correct answers), RL systematically selects moves that SFT almost never produced. This is creation of new preferences, not amplification—echoing AlphaZero's startling openings that violated centuries of human chess theory.
  • 3. Replication in mathematics

    The same laws hold for a 1B-parameter math model (pretraining on math text → SFT with reasoning traces → RL with verifiable answers):

  • Better-pretrained checkpoints reach higher post-RL math ability
  • RL curve slope scales roughly linearly with pretraining tokens
  • RL discovers solution paths SFT never attempted on hard problems
  • This suggests a cross-domain, domain-agnostic mapping between pretraining quality and RL outcomes.

    Implications for LLM training

  • Pretraining is the foundation: RL cannot compensate for insufficient pretraining; quality control in pretraining matters more than RL tuning.
  • Pretraining loss is an underrated metric: you can predict post-RL performance without running RL—valuable for checkpoint selection, early stopping, and compute allocation.
  • RL's value is discovery, not just optimization: reward functions focused on final-answer correctness (rather than imitating human reasoning formats) may unlock non-human reasoning paths; over-constraining chain-of-thought format may limit RL's exploratory power.
  • Cross-stage joint optimization: since pretraining loss encodes RL potential, future objectives could predictively maximize post-RL performance during pretraining itself.
  • Why chess works as a testbed

  • Fully deterministic rules, no ambiguity
  • Huge (~10^44 positions) but verifiable state space
  • Millennia of human knowledge for comparison
  • Automatic verification of game outcomes and puzzle answers
The authors call this the "Chess Testbed" and advocate controlled environments for studying AI fundamentals, invoking Feynman: if you can't explain it to a six-year-old, you don't really understand it.

References

1. Shen, J., Li, A., Rahman, S., et al. (2026). Understanding Reasoning from Pretraining to Post-Training. *arXiv:2607.16097* 2. Silver, D., et al. (2017). Mastering the Game of Go Without Human Knowledge. *Nature*, 550(7676), 354-359 3. Silver, D., et al. (2018). A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play. *Science*, 362(6419), 1140-1144 4. Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback. *NeurIPS 2022* 5. Guo, D., et al. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. *arXiv:2501.12948* 6. Kaplan, J., et al. (2020). Scaling Laws for Neural Language Models. *arXiv:2001.08361* 7. Hoffmann, J., et al. (2022). Training Compute-Optimal Large Language Models. *NeurIPS 2022* 8. Wei, J., et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. *NeurIPS 2022*

Tags

#llm-training#reinforcement-learning#pretraining#chess-testbed#reasoning#scaling-laws#deepseek-r1#verifiable-rewards

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446963