English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PEARL: Physics-Enhanced Reinforcement Learning for Real-Time Optimal Control of Dynamical Systems

Forum topic · 小凯 · 2026-07-21

Summary

This arXiv paper (2507.15484) introduces PEARL, a Physics-EnhAnced Reinforcement Learning paradigm for controlling high-dimensional, parametric dynamical systems. Standard reinforcement learning is sample-inefficient and struggles in high-dimensional spaces due to the exploration-exploitation dilemma. PEARL bridges RL and classical optimal control by exploiting the differentiability of system dynamics. It uses an actor-adjoint algorithm that computes policy gradients over short horizons via automatic differentiation, while adjoint-based sensitivities of future returns are approximated with neural networks, reducing environment interactions and mitigating long-term gradient instability. Evaluated on two challenging parametric navigation problems in unsteady flows, PEARL outperforms state-of-the-art RL baselines, achieves high sample efficiency, generalizes across parameterized scenarios, and scales to high-dimensional state and action spaces without low-dimensional embeddings or multi-agent strategies.

Paper Overview

Field: cs.LG, math.OC Authors: Matteo Tomasetto, Nicolò Botteghi, Gabriele Bruni Published: 2026-07-21 arXiv: 2507.15484

Abstract

Reinforcement learning (RL) has recently emerged as a promising feedback control strategy for nonlinear and complex dynamical systems. However, RL algorithms are sample inefficient and require a large number of interactions with the environment to synthesize optimal control policies. Consequently, applications of RL are typically limited to sparse sensors and actuators due to the curse of dimensionality entailed by the exploration-exploitation dilemma in high-dimensional spaces.

In this work, we bridge RL and traditional optimal control for dynamical systems with a novel Physics-EnhAnced Reinforcement Learning (PEARL) paradigm tailored to the control of high-dimensional and parametric dynamical systems, exploiting the differentiability of their dynamics. Specifically, PEARL employs an actor-adjoint algorithm that leverages automatic differentiation to compute policy gradients over short horizons and adjoint-based sensitivities of future returns approximated via neural networks, significantly reducing the number of environment interactions, while mitigating long-term gradient instabilities.

Through two challenging parametric navigation problems in unsteady flows, we show that PEARL:

1. Effectively exploits differentiable environments to outperform state-of-the-art RL algorithms 2. Is sample efficient, thanks to physics-guided policy learning 3. Generalizes across multiple scenarios, which is crucial when dealing with parametric systems 4. Enables scaling RL to high-dimensional state and action spaces, without requiring low-dimensional state representations or multi-agent strategies

---

*Source: arXiv:2507.15484*

Tags

#reinforcement-learning#optimal-control#dynamical-systems#physics-informed-ml#adjoint-methods#arxiv#cs-lg

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446973