[论文] Which Objectives Need a Dial? Predicting Objective Conflict and Coveri...
研究领域: NLP 作者: David Tsoi, Esra Dönmez 发布时间: 2026-09-25 arXiv: 2609.26929
论文概要
研究领域: NLP 作者: David Tsoi, Esra Dönmez 发布时间: 2026-09-25 arXiv: 2609.26929
中文摘要
人们持有多元、有时相互冲突的价值观,任何单一对齐模型都无法满足所有人。多元对齐因此需要可转向的模型,能以不同方式平衡相互竞争的目标。多目标直接偏好优化(MODPO)通过目标权重张成权衡的连续谱。本文研究两个问题:一个模型何时能同时改进两个目标?如何在不训练多个模型的情况下覆盖多种权衡?在来自 HelpSteer 与 UltraFeedback 的七对目标上,两项预训练测量能预测目标在人类标注数据上是对齐还是冲突,但在 AI 标注数据上不行——后者中回复长度与重复度混淆了奖励模型分数。在更广的权衡覆盖上,选择最近的已训练模型与合并模型参数都有帮助,但都无法稳定匹敌直接训练。这些发现为构建服务多元偏好的可转向模型提供了实用指导。
原文摘要
People hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective weight to span a continuum of trade-offs. We study two questions: when can one model improve two objectives simultaneously, and how can many trade-offs be covered without training a separate model for each? Across seven objective pairs from HelpSteer and UltraFeedback, two pre-training measurements predict whether objectives align or conflict for human-annotated data, but not for AI-annotated data, where response length and repetition confound reward-model scores. For broader trade-off coverage, se...
*自动采集于 2026-09-25*
#论文 #arXiv #NLP #小凯