[论文] A Unified Bellman Operator for Safety-Critical Reinforcement Learning

研究领域: ML 作者: Nishanth Arun Rao, Royina Karegoudra Jayanth, Benjamin Eysenbach, Jaime Fernández Fisac 发布时间: 2026-10-08 arXiv: 2610.12420

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Nishanth Arun Rao, Royina Karegoudra Jayanth, Benjamin Eysenbach, Jaime Fernández Fisac 发布时间: 2026-10-08 arXiv: 2610.12420

中文摘要

安全关键领域的强化学习需要在严格遵守安全约束的同时最大化任务性能。现有安全强化学习范式通常迫使人做出权衡:要么需要先验知识来提供严格的安全保证(如安全滤波器),要么支持联合学习但仅在平均意义上满足安全约束。本文提出一种新颖的贝尔曼算子,将性能和安全目标统一到一个联合值函数中。我们证明,使用该联合贝尔曼算子的时间差分学习在双时间尺度随机逼近框架下收敛。在快时间尺度上估计学习联合策略的安全值,在慢时间尺度上估计联合值。通过将极限动力学表述为占用平均微分包含并证明其渐近收敛到一组极限最优安全约束任务值函数来保证收敛。理论上,一旦收敛,所得最优策略在始终保持安全的同时最大化任务回报。在连续控制任务上的神经逼近实证评估展示了稳定收敛,测试时安全违规接近于零。

原文摘要

Reinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms typically force a trade-off: they either require a priori knowledge to provide strict safety guarantees (e.g., safety filters), or they enable joint learning but only satisfy safety constraints on average. In this work, we propose a novel Bellman operator that unifies performance and safety objectives into a joint value function. We show that temporal difference learning with the joint Bellman operator converges under a two-timescale stochastic approximation framework. On the fast timescale, the safety value of the learning joint policy is estimated, while the joint value is estimated on the slow timescale. Conv...


*自动采集于 2026-10-11*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens