[论文] Reinformed Dreamer: An Asymmetric World Model Efficiently Trained thro...
论文概要
研究领域: ML 作者: Gaspard Lambrechts, Adrien Bolland, Daniel Ebi, Damien Ernst 发布时间: 2026-07-28 arXiv: 2607.26040
中文摘要
就像人类在学习中受益于指导一样,强化学习算法也可能从奖励之外的额外监督中受益。在训练期间利用额外信息来学习更好的表征和行为一直是不对称强化学习的焦点。这种学习范式在部分可观测性下(当有额外状态信息可用时)已被证明有效,但在完全可观测性下(当有更精细的状态信息可用时)也同样有效。聚焦于基于模型的强化学习,我们研究了不对称学习对观测表征和特权信息表征的影响。首先,我们识别了已知的不对称模型算法Informed Dreamer在特权信息表征学习中的一个局限性。然后,我们提出了一种使用潜在引导的新型不对称表征学习目标,产生了一种名为Reinformed Dreamer的新算法。跨多个基准测试的实验表明,相比Dreamer,它比先前的不对称方法有更一致的改进。
原文摘要
Much like humans benefit from guidance while learning, reinforcement learning algorithms may benefit from additional supervision beyond rewards. Leveraging additional information during training to learn better representations and behaviors has been the focus of asymmetric reinforcement learning. This learning paradigm has proven effective under partial observability when additional state information is available, but also under full observability when more refined state information is available. Focusing on model-based reinforcement learning, we study the effect of asymmetric learning on observation representations and on privileged information representations. First, we identify a limitation in the privileged information representations learned by an asymmetric model-based algorithm know...
--- *自动采集于 2026-07-30*
#论文 #arXiv #ML #小凯
🌟 智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。
🎁 领取 2000万 Tokens