Paper Overview
- Field: NLP
- Authors: Sarah Wyer, Sue Black, Noura Al Moubayed
- Published: 2026-09-17
- arXiv: 2609.20779
Abstract (translated)
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this harm laundering.
Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not.
The pattern is most visible at GPT-5: Topic 5 (1,997 documents) frames breast cancer as a men's rights debate, while women-directed output shows zero equivalent topic clusters. Three independent classifiers rate this content as non-toxic. Sentiment scores invert at GPT-4: earlier models disparage women; later models over-correct.
At the GPT-4 alignment boundary, topic diversity for women-directed completions declines 36% relative to men (W/M = 0.58, versus 0.91 at GPT-2). REGARD representational harm correlates with release date (ρ = +0.55, p = .034), while Detoxify does not (ρ = -0.23, p = .42): toxicity scores fall as representational harm grows.
The authors formalize harm laundering as a three-criterion test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, reduced toxicity scores are not a sufficient proxy for reduced harm.
---
*Auto-collected on 2026-09-19*