论文概要
研究领域: ML
作者: Christopher M. Stewart, Preston Botter, Natalie Sarabosing, Muye Zhang, Rachel Phinnemore, Shalini Ghosh, Hong Shen, Hoda Heidari
发布时间: 2026-10-08
arXiv: 2610.12409
中文摘要
安全基准通常为一组数据集报告一个总分,每个数据集可能针对一个或多个安全相关属性,因此总分相近的模型可能有截然不同的属性特征。在单个属性层面比较模型更可行,但即使单个数据集的分数是否分离了任何单一属性往往也不清楚。「有害拒绝」——模型拒绝危险或违反策略提示的倾向——是这种属性的一个合理候选。我们检验它在 HELM Safety 中是否构成一个单一的、可测量的属性。使用构念效度框架(规定属性必须先于测试存在),我们从 HELM Safety 的四个可能针对有害拒绝的数据集开始,但发现三个已饱和。我们对剩余的数据集 HarmBench 进行两项心理测量检验,以确定像有害拒绝这样的单一属性是否可能支撑其分数。首先,多维项目反应理论模型强烈表明 HarmBench 并未测量单一属性。其次,差异项目功能分析发现,来自不同开发者但拒绝能力相同的模型在某些项目上得分不同。这些警示在特定领域匹配下基本消失,这一模式与聚合效应一致但不足以排除特定领域的开发者差异。放大来看,HarmBench 将不同的伤害行为折叠为一个分数,HELM 安全总分进一步将 HarmBench 和其他数据集的分数折叠为一个顶层数字。任何对数据集和项目取平均的安全分数都可能以这种方式隐藏饱和并混淆行为。我们主张,在用某个分数比较模型之前,它应先证明自己配得上单一属性的解读。
原文摘要
Safety benchmarks typically report one overall score for a suite of datasets, each of which may target one or more safety-related attributes, so models with similar overall scores can have very different attribute profiles. Comparing models is more tractable at the level of individual attributes, yet it is often unclear whether even a single dataset's scores isolate any single attribute. One plausible candidate for such an attribute is harmful refusal, a model's tendency to refuse dangerous or policy-violating prompts. We examine whether it constitutes a single, measurable attribute in HELM Safety. Using a construct validity framework that stipulates that an attribute must exist before a test can measure it, we start with HELM Safety's four datasets that might plausibly target harmful refu...
自动采集于 2026-10-11
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。