论文概要
研究领域: ML
作者: Zeyang Li, Sunbochen Tang, Navid Azizan
发布时间: 2026-09-14
arXiv: 2609.15915
中文摘要
元强化学习(meta-RL)使智能体能够仅凭有限经验适应未见任务。尽管前景广阔,meta-RL在真实任务中的应用受到安全要求的阻碍,而这一问题在先前的研究中尚未被充分探索。本文提出一个安全的meta-RL框架,在适应过程中显式考虑安全性。我们的关键洞察是在信息空间中推理安全性,信息空间同时捕捉物理状态和智能体对底层任务的信念。在该空间中,我们引入安全价值函数,衡量智能体无限期避开不安全区域的概率。我们证明该函数满足自洽条件和Bellman方程,使其可通过meta-RL学习。基于这一形式化,我们开发了一种安全meta-RL算法,学习安全价值函数并利用它进行安全过滤和约束策略优化。在meta-RL基准上的实验证明了所提方法的有效性。
原文摘要
Meta-reinforcement learning (meta-RL) enables agents to adapt to unseen tasks with limited experience. Despite its promise, the application of meta-RL in real-world tasks is hindered by safety requirements, which have been underexplored in prior work. In this paper, we propose a safe meta-RL framework that explicitly accounts for safety during adaptation. Our key insight is to reason about safety in the information space, which captures both the physical state and the agent's belief over the underlying task. Within this space, we introduce a safety value function that measures the probability of the agent avoiding unsafe regions indefinitely. We show that this function satisfies a self-consistency condition and a Bellman equation, which make it learnable via meta-RL. Based on this formulat...
自动采集于 2026-09-16
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。