English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Teacher Forcing as Generalized Bayes: Optimization Geometry Mismatch in Dynamical Systems Reconstruction

Forum topic · 小凯 · 2026-04-30

Summary

This arXiv paper (2504.21060) by Andre Herz, Daniel Durstewitz, Georgia Koppe, and colleagues analyzes why identity teacher forcing (ITF), despite enabling stable training of deterministic recurrent surrogates for chaotic dynamical systems, can be geometrically mismatched with the free-running model's marginal likelihood. Framing teacher forcing as an intervention-based prediction loss and thus a generalized Bayes update, the authors compare objective-induced curvatures of ITF and marginal likelihood in a probabilistic switching augmentation of almost-linear RNNs (AL-RNNs), estimating ambiguity-aware observed information via Louis' identity. They show that conditioning on a single forced regime path, as ITF does, inflates curvature, whereas marginal likelihood curvature decreases through a missing-information correction when multiple switching explanations remain plausible. In Lorenz-63 experiments, window-evidence fine-tuning improves held-out evidence but can degrade dynamical quantities of interest relative to ITF pretraining. The work clarifies optimization geometry trade-offs in dynamical systems reconstruction with recurrent neural networks.

Paper Overview

  • Field: Machine Learning
  • Authors: Andre Herz, Daniel Durstewitz, Georgia Koppe, et al.
  • arXiv: 2504.21060
  • Abstract (translated)

    Identity teacher forcing (ITF) enables stable training of deterministic recurrent surrogates for chaotic dynamical systems and has been highly effective for dynamical systems reconstruction (DSR) with recurrent neural networks (RNNs), including interpretable almost-linear RNNs (AL-RNNs). However, as an intervention-based prediction loss (and thus a generalized Bayes update), teacher forcing need not match the free-running model's marginal likelihood geometry.

    The authors compare the objective-induced curvatures of ITF and marginal likelihood in a probabilistic switching augmentation of AL-RNNs, estimating ambiguity-aware observed information via Louis' identity. In the switching setting studied, conditioning on a single forced regime path (as ITF does) inflates curvature, while marginal likelihood curvature is reduced through a missing-information correction when multiple switching explanations remain plausible.

    In Lorenz-63 experiments, window-evidence fine-tuning improves held-out evidence but may degrade dynamical quantities of interest (QoIs) compared with ITF-pretrained models.

    Key Points

  • ITF is effective for stable training of recurrent surrogates of chaotic systems, but is a generalized Bayes update based on interventional prediction, not the model's free-running likelihood.
  • Using Louis' identity, the paper shows ITF conditioning on a single forced regime path inflates the observed-information curvature.
  • Marginal likelihood applies a missing-information correction when switching regimes are ambiguous, lowering curvature.
  • On Lorenz-63, evidence-based fine-tuning improves held-out likelihood yet can worsen dynamical quantities of interest relative to ITF pretraining.
*Auto-collected 2026-04-30*

Tags

#machine-learning#teacher-forcing#dynamical-systems-reconstruction#rnn#bayesian-inference#arxiv#lorenz-63

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618918