论文概要
研究领域: NLP
作者: Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin
发布时间: 2026-09-16
arXiv: 2609.19144
中文摘要
直接偏好对齐方法因计算与内存效率而被广泛用于将大语言模型(LLM)与人类偏好对齐。然而,似然位移(likelihood displacement)现象促使研究者寻找替代方式,从似然 margin 较小的偏好对中提取信息。本文提出并分析了基于比较偏好优化(ComPO)——一种基于比较预言机(comparison oracle)的零阶对齐方法。ComPO 从这些偏好对中提取方向性信息,而不直接在它们上优化可微的偏好损失。我们在平滑性、梯度稀疏性以及预言机与潜在目标之间的兼容性条件下,建立了其基本离线方案的收敛保证。我们进一步提出在线 ComPO:保留离线比较机制,并利用无标注的策略生成实现相对于参考策略的反向 KL 控制。沿偏好微调的理论脉络,在局部覆盖与分布内成对奖励准确性条件下,我们建立了基本约束方案的性能保证。在 Mistral、Llama、Gemma-2、Qwen3 和 Gemma-3 上的实验表明,该方法优于现有直接对齐方法(包括长度控制的胜率),成对级别的诊断证据亦与缓解似然位移的推断一致。
原文摘要
Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margins. In this paper, we propose and analyze Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles. ComPO extracts directional information from these pairs without directly optimizing a differentiable preference loss on them. We establish a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective. We further introduce online ComPO, which retains the offli...
自动采集于 2026-09-18
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。