Paper Overview
Field: Machine Learning Authors: Dominik Żurek, Kamil Faber, Marcin Pietron, et al. arXiv: 2504.21087
Abstract
Continual offline reinforcement learning (CORL) aims to learn a sequence of tasks from datasets collected over time while preserving performance on previously learned tasks. This setting corresponds to domains where new tasks arise over time, but adapting the model in live environment interactions is expensive, risky, or impossible. However, CORL inherits the dual difficulty of offline reinforcement learning and adapting while preventing catastrophic forgetting.
Replay-based continual learning approaches remain a strong baseline but incur memory overhead and suffer from distribution mismatch between replayed samples and newly learned policies. At the same time, architectural continual learning methods have shown strong potential in supervised learning but remain underexplored in CORL.
This paper proposes TSN-Affinity, a novel CORL method based on TinySubNetworks and the Decision Transformer. The method implements task-specific parameterization with controlled knowledge sharing through an RL-aware reuse policy, which routes tasks based on action compatibility and latent similarity.
Evaluation
The method is evaluated on:
- Atari game benchmarks (discrete control)
- Simulated Franka Emika Panda robotic arm manipulation tasks (continuous control)
- Sparse subnetworks demonstrate strong retention capabilities
- Routing further improves multi-task performance
- Similarity-guided architectural reuse is a strong and feasible alternative to replay-based strategies in CORL settings
Findings
*Source: arXiv:2504.21087*