English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Pretraining Data Can Be Poisoned through Computational Propaganda: Public Discussion Interfaces as an Attack Vector

Forum topic · 小凯 · 2026-07-18

Summary

A paper by Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, and Kyle Lo (arXiv:2607.15267) demonstrates that poisoning attacks on language model pretraining data are feasible beyond controlled settings via public discussion interfaces—an existing web-scale content injection mechanism. Prior poisoning research mainly targeted established sources like Wikipedia, which do not reflect the scale and heterogeneity of real pretraining corpora and ignored how poisoned data interacts with data curation pipelines. The authors introduce HalfLife, a novel analysis method for estimating whether adversarial content survives web crawling and data curation and is included in web-crawl-based LM training data. Using HalfLife, they explore the feasibility of poisoning pretraining corpora at web scale through open discussion platforms. The study highlights the importance of estimating adversarial content inclusion in pretraining data and establishes third-party web content as a plausible vector for attacking LM pretraining.

Paper Overview

Research Area: NLP

Authors: Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, Kyle Lo

Published: 2026-07-16

arXiv: 2607.15267

Abstract

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines.

This paper demonstrates that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces.

Key Contributions

  • Attack feasibility beyond controlled settings: Shows that web-scale content injection via public discussion interfaces can serve as a realistic pretraining data poisoning vector, unlike prior work limited to established sources such as Wikipedia.
  • HalfLife: A novel analysis method for estimating the inclusion of adversarial content in web-crawl-based LM training data, measuring whether malicious content survives web crawling and data curation pipelines.
  • Feasibility analysis: Uses HalfLife to explore the feasibility of poisoning pretraining corpora at web scale through open discussion interfaces.

Implications

The analysis demonstrates the importance of estimating whether poisoned injections make it into pretraining data, and establishes third-party web content as a plausible vector for attacking language model pretraining—underscoring the need for data curation pipelines that account for adversarial content.

---

*Auto-collected on 2026-07-18*

Tags

#nlp#llm#data-poisoning#pretraining#security#arxiv#halflife#data-curation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178433596