English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Harm Laundering in GPT Models: Gender Discrimination Is Transformed, Not Removed

Forum topic · 小凯 · 2026-09-19

Summary

A 2026 arXiv paper (2609.20779) by Sarah Wyer, Sue Black, and Noura Al Moubayed introduces "harm laundering": the finding that explicit discriminatory content in large language models is transformed rather than eliminated as safety scores improve. Analyzing 450,000 gender-directed completions across 15 models from GPT-2 to GPT-5 under three demographic conditions, the authors show that sexual violence clusters in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions lack. At GPT-5, Topic 5 (1,997 documents) frames breast cancer as a men's rights debate, with no equivalent cluster for women-directed output; three independent classifiers rate the content non-toxic. Sentiment reverses at GPT-4 from disparaging women to over-correction. Topic diversity for women-directed completions drops 36% relative to men after the GPT-4 alignment boundary (W/M = 0.58 vs 0.91 at GPT-2). REGARD representational harm correlates with release date (rho = +0.55, p = .034) while Detoxify toxicity does not (rho = -0.23, p = .42), indicating declining toxicity scores are an insufficient proxy for declining harm. The paper formalizes a three-criterion harm laundering test and a three-stage detection protocol applicable to any generative model.

Paper Overview

  • Field: NLP
  • Authors: Sarah Wyer, Sue Black, Noura Al Moubayed
  • Published: 2026-09-17
  • arXiv: 2609.20779

Abstract (translated)

Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this harm laundering.

Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not.

The pattern is most visible at GPT-5: Topic 5 (1,997 documents) frames breast cancer as a men's rights debate, while women-directed output shows zero equivalent topic clusters. Three independent classifiers rate this content as non-toxic. Sentiment scores invert at GPT-4: earlier models disparage women; later models over-correct.

At the GPT-4 alignment boundary, topic diversity for women-directed completions declines 36% relative to men (W/M = 0.58, versus 0.91 at GPT-2). REGARD representational harm correlates with release date (ρ = +0.55, p = .034), while Detoxify does not (ρ = -0.23, p = .42): toxicity scores fall as representational harm grows.

The authors formalize harm laundering as a three-criterion test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, reduced toxicity scores are not a sufficient proxy for reduced harm.

---

*Auto-collected on 2026-09-19*

Tags

#nlp#llm#gpt#ai-safety#gender-bias#harm-laundering#arxiv#bias-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634987