This post discusses Google DeepMind's paper "The Topological Trouble With Transformers" (Mozer et al., 2026, arXiv:2604.17121) and its implications for large language models and AI agents.
Key points
- Transformers are directed acyclic graphs (DAGs): information flows only upward (shallow → deep layers) and leftward (earlier → later positions). There is no recurrence or feedback.
- State tracking requires
s_t = f(s_{t-1}, x_t), meaning each new state must live deeper than the previous one. After N tokens, the state representation has been pushed to layer N — but the model only has L layers. This is depth exhaustion. - Chain-of-Thought (CoT) works by printing deep-layer representations into output tokens, which become shallow inputs next round — effectively moving signals back to shallow layers. It is a costly patch, not a fix.
- Layers 1–5: the embedding blends "river" and "money" senses
- Layer ~6: the model correctly converges on the river sense
- But this deep disambiguation is invisible to shallow layers of subsequent tokens, so when "ATM" appears, shallow layers fall back on shallow co-occurrence statistics (bank → ATM → money), giving the wrong answer.
- Depth recurrence (Looped Transformer, Universal Transformer, RINS): adds expressivity but does not solve state tracking — parallel propagation across positions still pushes state deeper.
- Step-level recurrence is what truly enables unbounded state tracking, but it prevents parallelized training. Compromises include:
- Mamba / SSMs: parallel training (convolution-like), RNN-like inference — but linear updates are no more expressive than a standard Transformer (Merrill et al., 2025).
- DeltaNet: fast weight programming + delta rule; with eigenvalue ranges extended to negatives (Grazzi et al., 2025), it trains in parallel and exceeds standard Transformer expressivity.
- RWKV-7: expressive dynamic state evolution.
- COCONUT / Hierarchical Reasoning Model: latent, multi-cycle-per-token reasoning.
- Don't worship long context: scaling from 128K to 1M tokens does not fix implicit state tracking.
- CoT is not free: the thinking tax accumulates over long-horizon agents.
- Architecture sets the ceiling: for long-range consistency, multi-hop reasoning, and dynamic environments, consider Mamba/RWKV-7/DeltaNet, external memory systems, or hybrid fast-feedforward + slow-recurrent designs.
- The end goal is a "continuous inner monologue": a model that combines Transformer-style parallel training with RNN-style evolving internal state.
- Mozer et al. (2026). *The Topological Trouble With Transformers.* Google DeepMind, arXiv:2604.17121.
- Merrill & Sabharwal (2025). *The Expressive Power of Transformers with Chain of Thought.* ICLR.
- Giannou et al. (2023). *Looped Transformers as Programmable Computers.* ICML.
- Gu & Dao (2024). *Mamba: Linear-Time Sequence Modeling with Selective State Spaces.* ICML.
- Schlag et al. (2021). *Linear Transformers Are Secretly Fast Weight Memory Systems.* NeurIPS.
- Peng et al. (2025). *RWKV-7: Expressive Dynamic State Evolution.*
- Hao et al. (2025). *COCONUT: Continuous Latent Thought.*
- Allen-Zhu & Li (2025). *Canon Layers: Aligning Representations Across Steps.*
Empirical examples from the post
Number-guessing game. Gemini 3 fails to maintain a simple [low, high] interval: after answering "lower" to a guess of 60, it answers "higher" to a guess of 70. Even Gemini 3 *Thinking* writes in its internal monologue that the hidden number is 42, yet still answers "lower" when the user guesses 42 — an architectural defect, not a hallucination.
The "bank" experiment. Using Patchscopes, DeepMind traced activations layer by layer for the ambiguous word "bank":
Why CoT is a "thinking tax"
| Cost | Description | Scaling | |---|---|---| | Compute | Generating explicit steps for automatic inferences | Linear in thought length | | Context | Thought tokens crowd out usable history | Reduces available context | | Latency | Sequential generation of reasoning tokens | Multiplicative slowdown | | Memory | KV-cache grows with thought tokens | Linear |
DeepMind's point: reasoning that humans do automatically and unconsciously (e.g., disambiguating a word) should not require laborious explicit chains. Moreover, not all state can be externalized into natural language.
A taxonomy of recurrent architectures
Two axes: recurrence axis (depth, step-level, or both) × token-to-cycle ratio (>1 parallel chunks, =1 standard, <1 latent thinking).
DeepMind's five research directions
1. Hybrid models combining gated linear attention / gated delta nets with standard Transformer blocks. 2. Approximating state tracking in feedforward Transformers via training objectives and structural priors (a stopgap). 3. Coarse-grained recurrence at sentence or semantic-chunk level (e.g., Sentence Gestalt). 4. Representational alignment enabled by residual connections (e.g., Canon Layers), allowing variable-depth models to work without full retraining. 5. Multi-stage training: parallel feedforward pretraining, then recurrent fine-tuning, with truncated gradients and backpropagation through recurrence.