[论文] Gradient-Aligned Pair Selection for Personalized Preference Optimizati...

研究领域: ML 作者: Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou 发布时间: 2026-10-05 arXiv: 2610.00061

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou 发布时间: 2026-10-05 arXiv: 2610.00061

中文摘要

个性化大语言模型需要对齐生成行为与用户特定偏好,而非聚合质量。直接偏好优化(DPO)为偏好学习提供稳定框架,但其在个性化场景下的成效关键取决于如何选择偏好对。现有方法多依赖启发式标准(如似然极端值),使优化与显式用户效用解耦,可能导致个性化退化。我们通过分析期望用户效用梯度与 DPO 更新方向的一阶交互,将个性化偏好学习形式化为几何对齐优化问题。分析揭示:离策略采样下,当偏好边距与效用梯度方向对齐时,DPO 更新从纯误差纠正信号转变为类强化更新。这一视角暴露偏好对选择是几何决策——决定偏好优化推进还是阻碍个性化。受此启发,我们提出 GAP-DPO(几何对齐偏好 DPO),迭代执行效用感知、几何对齐的偏好对选择,并经按轮次再生成控制分布漂移。个性化文本生成基准实验显示,GAP-DPO 在风格保真度、偏好对齐和生成质量上一致优于标准 DPO 变体。结果确立梯度对齐为个性化偏好优化的统一原则,证明偏好对选择是优化几何的内在组成,而非启发式预处理步骤。

原文摘要

Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference pairs are selected. Existing approaches typically rely on heuristic criteria, such as likelihood-based extremes, which decouple optimization from explicit user utility and can lead to degraded personalization. We formalize personalized preference learning as a geometry-aligned optimization problem by analyzing the first-order interaction between gradients of expected user utility and DPO update directions. Our analysis reveals that, under off-policy sampling, the DPO update tr...


*自动采集于 2026-10-05*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens