[论文] FERPO: Forward Entropy-Regularized Policy Optimization
研究领域: ML 作者: Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv 发布时间: 2026-10-01 arXiv: 2610.02198
论文概要
研究领域: ML 作者: Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv 发布时间: 2026-10-01 arXiv: 2610.02198
中文摘要
连续控制在线强化学习的若干最先进方法,利用学习得到的 critic 的动作梯度来改进策略。然而 critic 通常被训练来预测回报,精确的价值预测并不必然产生精确的动作导数,可能导致不可靠的策略更新。我们提出前向熵正则化策略优化(FERPO),一种 on-policy 最大熵强化学习算法:使用 critic 的值而不对 critic 关于动作求导,即可完成策略改进。FERPO 从由熵和 KL 散度正则化的策略改进目标中推导出最优目标动作分布,然后通过最小化前向 KL 目标将 actor 拟合到该分布,并用 rollout 策略采样的动作、经自归一化重要性采样(SNIS)估计。KL 正则化限制目标分布偏离 rollout 策略的程度,使重要性权重保持良好性质。与可能偏向目标分布部分模态的反向 KL 目标不同,前向 KL 目标鼓励覆盖多个高值模态,从而促进探索。MuJoCo Playground 与 ManiSkill 上的实验和消融显示了有竞争力的性能与样本效率增益;计算基准还表明其 actor 更新快于 REPPO。
原文摘要
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and KL divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimate...
*自动采集于 2026-10-04*
#论文 #arXiv #ML #小凯