Overview
Field: Machine Learning Authors: Tiberiu Musat, Tiago Pimentel, Nicholas Zucchet, Thomas Hofmann Published: 2026-07-13 arXiv: 2607.11875
A theoretical framework is presented to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have been mostly tied to specific tasks, this paper studies a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning.
Key points
- Invariant manifold: In this task class, the authors theoretically prove that the training dynamics of attention models can be confined to a highly interpretable, low-dimensional invariant manifold.
- Interpretable coordinates: On this manifold, learning dynamics are captured by a handful of interpretable coordinates rather than millions of parameters, making both theoretical and empirical analysis more tractable.
- In-context vs. in-weights learning: The framework characterizes how data statistics govern the competition between in-context learning and in-weights learning.
- Circuit selection: It examines how random initialization determines which circuit "wins" when multiple solutions exist.
- Circuit detection: The manifold-related coordinate framework can be used to automatically detect which circuits a trained model has learned.
Significance
Treating circuit formation as a low-dimensional dynamical phenomenon is a step toward a predictive theory of how Transformers learn.
Source: arXiv:2607.11875