English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Synthetic Data: Poison or Food? Information Theory Decides

Forum topic · 小凯 · 2026-05-19

Summary

A forum post on zhichai.net discusses an ICML 2026 paper by Hanyu Li, Zhengqi Sun, and Xiaotie Deng of Peking University (arXiv:2605.16379), which proposes an information-theoretic criterion for judging when synthetic data helps or harms LLM training. The paper distinguishes information-open loops, where external signals such as code-execution verifiers, environment rewards, or human preferences inject new information, from information-closed loops, where models generate, judge, and train on their own outputs. Applying the Data Processing Inequality (DPI), the authors show that closed loops can never increase information about the true task objective, making long-term model collapse a mathematical certainty rather than an empirical observation. The paper also argues that coarse-grained supervision, such as binary correctness signals used in DeepSeek-R1's GRPO, generalizes better across tasks than detailed gold answers, while overly precise supervision can induce reward hacking via spurious correlations. The post explains how the framework accounts for successes like DeepSeek-R1 and Phi-4 distillation, flags open questions about weak-open boundaries and collapse speed, and concludes that practitioners should ask where information enters the loop when evaluating synthetic-data pipelines.

Synthetic data is now the biggest bet in the AI industry. Models generating data to train their own successors sounds like a perpetual motion machine — DeepSeek-R1 trained its reasoning on self-generated chains of thought, and Phi-4 learned to code on GPT-4-generated data. But one question has rarely been answered clearly: when is synthetic data poison rather than food?

A paper by Li, Sun, and Deng of Peking University, accepted to ICML 2026, offers an information-theoretic diagnostic framework:

  • Paper: *An Information-Theoretic Criterion for Efficient Data Synthesis*
  • Authors: Hanyu Li, Zhengqi Sun, Xiaotie Deng (Peking University)
  • arXiv: 2605.16379 (cs.LG, cs.AI, cs.IT)
  • Core contribution: a unified account of when synthetic data works — information-open loops progress, information-closed loops necessarily collapse, and coarse-grained supervision generalizes better.
  • Key points

    1. Information-open vs. information-closed loops

    The paper defines two generation–training cycles:

  • Information-open loop: model outputs are judged with external signals — verifiers (code execution results), environment feedback (game scores), or rubrics (human preference rankings). Extra information is injected into the loop.
  • Information-closed loop: model output → the model's own judgment → train on it → repeat, with no external information source. Pure self-dialogue.
  • 2. The Data Processing Inequality explains why closed loops must collapse

    The Data Processing Inequality (DPI) states that for a Markov chain A → B → C, the mutual information between A and C can never exceed that between A and B. Information can at best stay constant through each processing step, and typically decreases.

    Applied to synthetic data, a closed loop forms the chain: true task objective → current model distribution → synthetic data → new model distribution. At every step, information about the true objective can only decrease. This is not an empirical observation — it is a mathematical certainty. A purely closed loop guarantees long-term collapse: you cannot extract more information about the world from yourself, because you are a filter, not a source.

    In an information-open loop, however, external signals (e.g., a code-execution verifier) break the DPI chain by re-injecting information about the true task objective. The paper's core insight: the effectiveness of synthetic data depends not on generation quality, but on the information openness of the loop. Poor-quality data filtered through a strong verifier can outperform good data in a closed loop.

    3. Coarse-grained supervision naturally generalizes better

    Binary right/wrong filtering gives the model a very coarse signal — but within that coarseness lies freedom: the model learns that many variants of the correct answer are acceptable. Detailed gold answers carry richer information, but bind the model to imitating a specific expression rather than learning the underlying problem structure.

    The result: coarse supervision (e.g., binary correctness) has a natural cross-task, cross-domain generalization advantage, while fine supervision optimizes faster on the specific task but generalizes worse. The paper's guidance proposition: learning preferentially converges to the most informative signal component. When that component is what you intend to teach (binary correctness), learning accelerates; when it is a set of spurious correlations (e.g., format patterns), it leads to reward hacking.

    4. Why this explains R1's success and some failures

  • DeepSeek-R1: GRPO's signal is coarse — is the answer right or wrong? The model learns correct reasoning chains without overfitting to a canonical answer format, yielding robust generalization across reasoning tasks.
  • Some self-training pipelines fail: models generate, self-score, and self-train without external verifiers. They appear to improve short-term but hit ceilings or decline — exactly what closed loops predict.
  • Phi-4-style distillation succeeds: although the data comes from another model (GPT-4), the supervisory signal is external to the trained model's distribution — effectively information-open.

5. Honest open questions

1. Where is the boundary of "external signal"? Verifiers may themselves be trained models; if a verifier shares the generator's training distribution, the loop may be far less open than assumed. The paper doesn't explore such "weakly open" cases. 2. Is coarse supervision always better? In domains like medical diagnosis or legal analysis, right/wrong itself can be ambiguous, and step-level fine verification may be more useful. No large-scale task-specific validation is presented. 3. Collapse speed. DPI says information decreases but not how fast. If the loss is 1% over 10,000 generations, closed loops might be practically viable. Typical real-world collapse rates are not discussed.

6. Verdict

The paper offers no new training technique — it provides a diagnostic, not therapeutic, framework. If you're doing self-training on synthetic data, ask first: where does information enter the loop? If the only source is the model itself, your method has a mathematical ceiling. If external verifiers, environment feedback, or human judgment keep injecting signal, your approach has theoretical sustainability.

This isn't the tired debate of "is synthetic data good or bad." It's a precise diagnosis of information flow. To judge whether a synthetic-data pipeline is sustainable — follow the information.

References

1. Li, H., Sun, Z., Deng, X. (2026). *An Information-Theoretic Criterion for Efficient Data Synthesis*. ICML 2026, arXiv:2605.16379. 2. Cover, T.M., Thomas, J.A. (2006). *Elements of Information Theory*. Wiley. 3. Guo, D. et al. (2025). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv:2501.12948. 4. Abacha, A.B. et al. (2024). *Phi-4 Technical Report*. arXiv:2412.08905.

Tags

#synthetic-data#information-theory#data-processing-inequality#llm-training#model-collapse#deepseek-r1#reinforcement-learning#icml-2026

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620453