Synthetic data is now the biggest bet in the AI industry. Models generating data to train their own successors sounds like a perpetual motion machine — DeepSeek-R1 trained its reasoning on self-generated chains of thought, and Phi-4 learned to code on GPT-4-generated data. But one question has rarely been answered clearly: when is synthetic data poison rather than food?
A paper by Li, Sun, and Deng of Peking University, accepted to ICML 2026, offers an information-theoretic diagnostic framework:
- Paper: *An Information-Theoretic Criterion for Efficient Data Synthesis*
- Authors: Hanyu Li, Zhengqi Sun, Xiaotie Deng (Peking University)
- arXiv: 2605.16379 (cs.LG, cs.AI, cs.IT)
- Core contribution: a unified account of when synthetic data works — information-open loops progress, information-closed loops necessarily collapse, and coarse-grained supervision generalizes better.
- Information-open loop: model outputs are judged with external signals — verifiers (code execution results), environment feedback (game scores), or rubrics (human preference rankings). Extra information is injected into the loop.
- Information-closed loop: model output → the model's own judgment → train on it → repeat, with no external information source. Pure self-dialogue.
- DeepSeek-R1: GRPO's signal is coarse — is the answer right or wrong? The model learns correct reasoning chains without overfitting to a canonical answer format, yielding robust generalization across reasoning tasks.
- Some self-training pipelines fail: models generate, self-score, and self-train without external verifiers. They appear to improve short-term but hit ceilings or decline — exactly what closed loops predict.
- Phi-4-style distillation succeeds: although the data comes from another model (GPT-4), the supervisory signal is external to the trained model's distribution — effectively information-open.
Key points
1. Information-open vs. information-closed loops
The paper defines two generation–training cycles:
2. The Data Processing Inequality explains why closed loops must collapse
The Data Processing Inequality (DPI) states that for a Markov chain A → B → C, the mutual information between A and C can never exceed that between A and B. Information can at best stay constant through each processing step, and typically decreases.
Applied to synthetic data, a closed loop forms the chain: true task objective → current model distribution → synthetic data → new model distribution. At every step, information about the true objective can only decrease. This is not an empirical observation — it is a mathematical certainty. A purely closed loop guarantees long-term collapse: you cannot extract more information about the world from yourself, because you are a filter, not a source.
In an information-open loop, however, external signals (e.g., a code-execution verifier) break the DPI chain by re-injecting information about the true task objective. The paper's core insight: the effectiveness of synthetic data depends not on generation quality, but on the information openness of the loop. Poor-quality data filtered through a strong verifier can outperform good data in a closed loop.
3. Coarse-grained supervision naturally generalizes better
Binary right/wrong filtering gives the model a very coarse signal — but within that coarseness lies freedom: the model learns that many variants of the correct answer are acceptable. Detailed gold answers carry richer information, but bind the model to imitating a specific expression rather than learning the underlying problem structure.
The result: coarse supervision (e.g., binary correctness) has a natural cross-task, cross-domain generalization advantage, while fine supervision optimizes faster on the specific task but generalizes worse. The paper's guidance proposition: learning preferentially converges to the most informative signal component. When that component is what you intend to teach (binary correctness), learning accelerates; when it is a set of spurious correlations (e.g., format patterns), it leads to reward hacking.
4. Why this explains R1's success and some failures
5. Honest open questions
1. Where is the boundary of "external signal"? Verifiers may themselves be trained models; if a verifier shares the generator's training distribution, the loop may be far less open than assumed. The paper doesn't explore such "weakly open" cases. 2. Is coarse supervision always better? In domains like medical diagnosis or legal analysis, right/wrong itself can be ambiguous, and step-level fine verification may be more useful. No large-scale task-specific validation is presented. 3. Collapse speed. DPI says information decreases but not how fast. If the loss is 1% over 10,000 generations, closed loops might be practically viable. Typical real-world collapse rates are not discussed.
6. Verdict
The paper offers no new training technique — it provides a diagnostic, not therapeutic, framework. If you're doing self-training on synthetic data, ask first: where does information enter the loop? If the only source is the model itself, your method has a mathematical ceiling. If external verifiers, environment feedback, or human judgment keep injecting signal, your approach has theoretical sustainability.
This isn't the tired debate of "is synthetic data good or bad." It's a precise diagnosis of information flow. To judge whether a synthetic-data pipeline is sustainable — follow the information.
References
1. Li, H., Sun, Z., Deng, X. (2026). *An Information-Theoretic Criterion for Efficient Data Synthesis*. ICML 2026, arXiv:2605.16379. 2. Cover, T.M., Thomas, J.A. (2006). *Elements of Information Theory*. Wiley. 3. Guo, D. et al. (2025). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv:2501.12948. 4. Abacha, A.B. et al. (2024). *Phi-4 Technical Report*. arXiv:2412.08905.