[论文] Predicting Alignment Generalization with Value Representations

研究领域: NLP 作者: Andy Liu, Mehar Bhatia, Karolina Stanczak, Mona Diab, Vered Shwartz, Daniel Fried 发布时间: 2026-10-08 arXiv: 2610.12410

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: NLP 作者: Andy Liu, Mehar Bhatia, Karolina Stanczak, Mona Diab, Vered Shwartz, Daniel Fried 发布时间: 2026-10-08 arXiv: 2610.12410

中文摘要

LLM 开发者对模型进行后训练,使其表现出亲社会价值观和行为特征,这些都在对齐目标中被枚举。然而,尽管近期的后训练进展使模型在对齐评估中得分很高,但在狭窄行为集上训练模型仍会以其行为在意外的方向上影响模型在未见过情境和环境中的表现。本文建立了对齐泛化预测任务——即预测在给定价值上微调模型将如何改变其在广泛留出价值上的行为。我们对现代对齐目标中 66 种价值的对齐泛化效应进行了大规模分析,并在对齐泛化预测任务上基准测试了多种表征技术。我们发现,基于模型在上下文中应用价值时的激活的表征,显著优于基于价值文本描述的方法。具体而言,最佳基于激活的方法与泛化矩阵的相关性达 0.45,而基于描述的基线仅为 0.05。然后我们展示了预测对齐泛化的表征在下游任务中的适用性:用它们衡量多价值对齐目标中价值的相似性,发现与模型鲁棒性显著相关。最后,我们展示了共享的、模型无关的价值空间的初步证据,并用它开发了首个基于经验泛化动态的 LLM 价值分类法。我们的工作展示了对 LLM 价值泛化研究的重要性及其在更经验化的模型行为设计和训练中的应用。

原文摘要

LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values. We conduct a large-scale analysis of alignment generalization effects across 66 values found in modern alignment targets, and benchmark representational techniques on the alignment generalization prediction task. ...


*自动采集于 2026-10-11*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens