Paper Overview
Field: Machine Learning Authors: Austin Coursey, Abel Diaz-Gonzalez, Marcos Quinones-Grueiro, Gautam Biswas Published: 2026-04-21 arXiv: 2604.19737
Abstract
Reinforcement learning (RL) offers a compelling data-driven paradigm for synthesizing controllers for complex systems when accurate physical models are unavailable; however, most existing control-oriented RL methods assume stationarity and, therefore, struggle in real-world non-stationary deployments where system dynamics and operating conditions can change unexpectedly. Moreover, RL controllers acting in physical environments must satisfy safety constraints throughout their learning and execution phases, rendering transient violations during adaptation unacceptable.
Although continual RL and safe RL have each addressed non-stationarity and safety, respectively, their intersection remains comparatively unexplored, motivating the study of safe continual RL algorithms that can adapt over a system's lifetime while remaining safe.
Key Contributions
- Benchmark environments: Introduces three environments capturing safety-critical continual adaptation.
- Systematic evaluation: Benchmarks representative methods from safe RL, continual RL, and their combination.
- Empirical findings: Reveals a fundamental tension between maintaining safety constraints and preventing catastrophic forgetting under non-stationary dynamics; existing methods typically fail to achieve both objectives simultaneously.
- Regularization-based approaches: Examined as a partial mitigation of this trade-off, with benefits and limitations characterized.
- Outlook: Outlines open challenges and research directions for safe, resilient learning-based controllers that operate autonomously in changing environments.
*Auto-collected on 2026-04-23*