Paper Overview
- Field: Machine Learning
- Authors: Andre Herz, Daniel Durstewitz, Georgia Koppe, et al.
- arXiv: 2504.21060
- ITF is effective for stable training of recurrent surrogates of chaotic systems, but is a generalized Bayes update based on interventional prediction, not the model's free-running likelihood.
- Using Louis' identity, the paper shows ITF conditioning on a single forced regime path inflates the observed-information curvature.
- Marginal likelihood applies a missing-information correction when switching regimes are ambiguous, lowering curvature.
- On Lorenz-63, evidence-based fine-tuning improves held-out likelihood yet can worsen dynamical quantities of interest relative to ITF pretraining.
Abstract (translated)
Identity teacher forcing (ITF) enables stable training of deterministic recurrent surrogates for chaotic dynamical systems and has been highly effective for dynamical systems reconstruction (DSR) with recurrent neural networks (RNNs), including interpretable almost-linear RNNs (AL-RNNs). However, as an intervention-based prediction loss (and thus a generalized Bayes update), teacher forcing need not match the free-running model's marginal likelihood geometry.
The authors compare the objective-induced curvatures of ITF and marginal likelihood in a probabilistic switching augmentation of AL-RNNs, estimating ambiguity-aware observed information via Louis' identity. In the switching setting studied, conditioning on a single forced regime path (as ITF does) inflates curvature, while marginal likelihood curvature is reduced through a missing-information correction when multiple switching explanations remain plausible.
In Lorenz-63 experiments, window-evidence fine-tuning improves held-out evidence but may degrade dynamical quantities of interest (QoIs) compared with ITF-pretrained models.