论文概要
研究领域: NLP
作者: Sarah Wyer, Sue Black, Noura Al Moubayed
发布时间: 2026-09-17
arXiv: 2609.20779
中文摘要
大语言模型的安全评估依赖于表层形式分类器,这些分类器报告模型代际间危害分数的下降。我们提供的证据表明,这一方法论存在系统性的不完整:显性歧视内容被转化而非消除。我们称之为"危害洗白"(harm laundering)。通过分析跨越 GPT-2 至 GPT-5 的15个模型(OpenAI GPT 谱系;三种人口统计条件)的45万条性别定向补全,我们表明:GPT-2 女性定向输出中普遍的性暴力主题簇到 GPT-4 时消失,而男性定向补全获得了女性定向补全所没有的正面表征领域(照护、情感范围、盟友身份)。这一模式在 GPT-5 上最为明显:主题5(1,997份文档)将乳腺癌框架化为男性权利辩论,而女性定向输出中出现零等价主题簇。三个独立分类器将该内容评分为非毒性。情感分数在 GPT-4 处反转:早期模型贬低女性;后期模型过度矫正。在 GPT-4 对齐边界处,女性定向补全的主题多样性相对男性下降36%(W/M = 0.58,GPT-2 时为 0.91)。REGARD 表征危害差异与发布日期相关(ρ = +0.55, p = .034),而 Detoxify 不相关(ρ = -0.23, p = .42):毒性分数下降的同时,表征危害在增长。我们将危害洗白形式化为三标准检验,并提供适用于任何生成模型的三阶段检测协议。在 OpenAI GPT 谱系内,毒性分数降低并不是危害降低的充分代理指标。
原文摘要
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this harm laundering. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic 5 (1,997 documents) frames breast cancer as a...
自动采集于 2026-09-19
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。