论文概要
研究领域: ML
作者: Ondrej Bajgar, Peter Tisnikar, Alessandro Abate et al. (5 authors)
发布时间: 2026-08-17
arXiv: 2608.16888
中文摘要
安全有益的AI需要系统能按照人类偏好学习和行动,但手工明确指定这些偏好往往不可行。逆向强化学习(IRL)通过从专家行为推断偏好(以奖励函数表示)来解决这一挑战。我们提出基于Q值的变分IRL(QVIRL),一种新颖的贝叶斯IRL方法,通过主要学习最优Q值上的变分分布,从专家演示中恢复奖励的后验分布。与以往方法不同,QVIRL结合了可扩展性和不确定性量化,这对安全关键应用和主动学习都很重要。我们在多种任务上展示了QVIRL在学徒学习中的强劲表现,包括网格世界、月球着陆器、高速公路环境和两款ATARI游戏,同时使用静态专家数据和主动学习。这是首个展示从原始像素观测进行训练的贝叶斯IRL方法。
原文摘要
The development of safe and beneficial AI requires that systems can learn and act in accordance with human preferences. However, explicitly specifying these preferences by hand is often infeasible. Inverse reinforcement learning (IRL) addresses this challenge by inferring preferences, represented as reward functions, from expert behaviour. We introduce Q-based Variational IRL (QVIRL), a novel Bayesian IRL method that recovers a posterior distribution over rewards from expert demonstrations via primarily learning a variational distribution over optimal Q-values. Unlike previous approaches, QVIRL combines scalability with uncertainty quantification, important for safety-critical applications as well as active learning. We demonstrate QVIRL's strong performance in apprenticeship learning acro...
自动采集于 2026-08-19
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。