[论文] A Zeroth-Order Paradigm for LLM Preference Alignment

研究领域: NLP 作者: Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin 发布时间: 2026-09-16 arXiv: 2609.19144

论文概要

研究领域: NLP 作者: Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin 发布时间: 2026-09-16 arXiv: 2609.19144

中文摘要

直接偏好对齐方法因计算与内存效率而被广泛用于将大语言模型(LLM)与人类偏好对齐。然而,似然位移(likelihood displacement)现象促使研究者寻找替代方式,从似然 margin 较小的偏好对中提取信息。本文提出并分析了基于比较偏好优化(ComPO)——一种基于比较预言机(comparison oracle)的零阶对齐方法。ComPO 从这些偏好对中提取方向性信息,而不直接在它们上优化可微的偏好损失。我们在平滑性、梯度稀疏性以及预言机与潜在目标之间的兼容性条件下,建立了其基本离线方案的收敛保证。我们进一步提出在线 ComPO:保留离线比较机制,并利用无标注的策略生成实现相对于参考策略的反向 KL 控制。沿偏好微调的理论脉络,在局部覆盖与分布内成对奖励准确性条件下,我们建立了基本约束方案的性能保证。在 Mistral、Llama、Gemma-2、Qwen3 和 Gemma-3 上的实验表明,该方法优于现有直接对齐方法(包括长度控制的胜率),成对级别的诊断证据亦与缓解似然位移的推断一致。

原文摘要

Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margins. In this paper, we propose and analyze Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles. ComPO extracts directional information from these pairs without directly optimizing a differentiable preference loss on them. We establish a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective. We further introduce online ComPO, which retains the offli...


*自动采集于 2026-09-18*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens