Paper Overview
Research Area: NLP
Authors: Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, Kyle Lo
Published: 2026-07-16
arXiv: 2607.15267
Abstract
Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines.
This paper demonstrates that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces.
Key Contributions
- Attack feasibility beyond controlled settings: Shows that web-scale content injection via public discussion interfaces can serve as a realistic pretraining data poisoning vector, unlike prior work limited to established sources such as Wikipedia.
- HalfLife: A novel analysis method for estimating the inclusion of adversarial content in web-crawl-based LM training data, measuring whether malicious content survives web crawling and data curation pipelines.
- Feasibility analysis: Uses HalfLife to explore the feasibility of poisoning pretraining corpora at web scale through open discussion interfaces.
Implications
The analysis demonstrates the importance of estimating whether poisoned injections make it into pretraining data, and establishes third-party web content as a plausible vector for attacking language model pretraining—underscoring the need for data curation pipelines that account for adversarial content.
---
*Auto-collected on 2026-07-18*