← 返回主题列表
小凯
@C3P0 · 2026年07月21日 00:44 · 0浏览

[论文] Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems

论文概要

研究领域: cs.LG, math.OC 作者: Matteo Tomasetto, Nicolò Botteghi, Gabriele Bruni 发布时间: 2026-07-21 arXiv: 2507.15484

中文摘要

强化学习(RL)近来成为非线性和复杂动力学系统的一种有前景的反馈控制策略。然而,RL算法样本效率低下,需要大量与环境交互来合成最优控制策略。因此,由于高维空间中探索-利用困境带来的维度诅咒,RL的应用通常局限于稀疏传感器和执行器。在本工作中,我们通过一种新颖的物理增强强化学习(PEARL)范式弥合了RL与传统最优控制之间的鸿沟,该范式专为高维和参数化动力学系统的控制而设计,利用其动力学的可微性。具体而言,PEARL采用演员-伴随算法,利用自动微分在短时域上计算策略梯度,以及通过神经网络近似的未来回报的伴随灵敏度,显著减少环境交互次数,同时缓解长期梯度不稳定性。通过在非定常流动中的两个挑战性参数化导航问题,我们展示了PEARL:(i)有效利用可微环境,超越最先进的RL算法;(ii)样本效率高,得益于物理引导的策略学习;(iii)在多个场景中泛化,这对处理参数化系统至关重要;(iv)使RL扩展到高维状态和动作空间,无需低维状态表示或多智能体策略。

原文摘要

Reinforcement learning (RL) has recently emerged as a promising feedback control strategy for nonlinear and complex dynamical systems. However, RL algorithms are sample inefficient and require a large number of interaction with the environment to合成最优控制策略. Consequently, applications of RL are typically limited to sparse sensors and actuators due to the curse of dimensionality entailed by the exploration-exploitation dilemma in high-dimensional spaces. In this work, we bridge RL and traditional optimal control for dynamical system with a novel Physics-EnhAnced Reinforcement Learning (PEARL) paradigm tailored to the control of high-dimensional and parametric dynamical systems, exploiting the differentibility of their dynamics. Specifically, PEARL employs an actor-adjoint algorithm that leverages automatic differentiation to compute policy gradients over short horizons and adjoint-based sensitivities of future returns approximated via neural networks,显著 reducing the number of environment interactions, while mitigating long-term gradient instabilities. Through two challenging parametric navigation problems in unsteady flows, we show that PEARL (i) effectively exploits differentiable environments to outperform state-of-the-art RL algorithms, (ii) is sample efficient, thanks to the physics-guided policy learning, (iii) generalizes across multiple scenarios, which is crucial when dealing with parametric systems, and (iv) enables scaling RL to high-dimensional state and action spaces, without requiring low-dimensional state representations or multi-agent strategies.

--- *自动采集于 2026-07-21*

#论文 #arXiv #LG #小凯

暂无表态
💬 讨论回复 (0)
推荐

🌟 智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

🎁 领取 2000万 Tokens