English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DeepMind: The Topological Trouble With Transformers — Why Bigger Context Windows Won't Fix LLMs

Forum topic · 小凯 · 2026-06-24

Summary

A Chinese tech forum post analyzes Google DeepMind's paper 'The Topological Trouble With Transformers' (Mozer et al., arXiv:2604.17121), arguing that the Transformer's directed acyclic graph (DAG) feedforward architecture fundamentally limits state tracking. Simple tasks like a number-guessing game expose Gemini 3's inability to maintain a [low, high] interval, and Patchscopes probing shows that word-sense disambiguation computed in deep layers becomes invisible to shallow layers of later tokens. Because state must flow to ever-deeper layers, models with L layers fail after N > L steps — 'depth exhaustion.' Chain-of-Thought is characterized as an expensive patch that externalizes latent state into tokens, incurring compute, latency, memory, and context costs. The post surveys DeepMind's taxonomy of recurrent architectures along depth/step axes and token-to-cycle ratios, highlighting step-level recurrence (Mamba, DeltaNet, RWKV-7) as promising, and outlines five research directions, concluding that agentic AI needs evolving internal state rather than longer context or explicit reasoning chains.

This post discusses Google DeepMind's paper "The Topological Trouble With Transformers" (Mozer et al., 2026, arXiv:2604.17121) and its implications for large language models and AI agents.

Key points

  • Transformers are directed acyclic graphs (DAGs): information flows only upward (shallow → deep layers) and leftward (earlier → later positions). There is no recurrence or feedback.
  • State tracking requires s_t = f(s_{t-1}, x_t), meaning each new state must live deeper than the previous one. After N tokens, the state representation has been pushed to layer N — but the model only has L layers. This is depth exhaustion.
  • Chain-of-Thought (CoT) works by printing deep-layer representations into output tokens, which become shallow inputs next round — effectively moving signals back to shallow layers. It is a costly patch, not a fix.
  • Empirical examples from the post

    Number-guessing game. Gemini 3 fails to maintain a simple [low, high] interval: after answering "lower" to a guess of 60, it answers "higher" to a guess of 70. Even Gemini 3 *Thinking* writes in its internal monologue that the hidden number is 42, yet still answers "lower" when the user guesses 42 — an architectural defect, not a hallucination.

    The "bank" experiment. Using Patchscopes, DeepMind traced activations layer by layer for the ambiguous word "bank":

  • Layers 1–5: the embedding blends "river" and "money" senses
  • Layer ~6: the model correctly converges on the river sense
  • But this deep disambiguation is invisible to shallow layers of subsequent tokens, so when "ATM" appears, shallow layers fall back on shallow co-occurrence statistics (bank → ATM → money), giving the wrong answer.
  • Why CoT is a "thinking tax"

    | Cost | Description | Scaling | |---|---|---| | Compute | Generating explicit steps for automatic inferences | Linear in thought length | | Context | Thought tokens crowd out usable history | Reduces available context | | Latency | Sequential generation of reasoning tokens | Multiplicative slowdown | | Memory | KV-cache grows with thought tokens | Linear |

    DeepMind's point: reasoning that humans do automatically and unconsciously (e.g., disambiguating a word) should not require laborious explicit chains. Moreover, not all state can be externalized into natural language.

    A taxonomy of recurrent architectures

    Two axes: recurrence axis (depth, step-level, or both) × token-to-cycle ratio (>1 parallel chunks, =1 standard, <1 latent thinking).

  • Depth recurrence (Looped Transformer, Universal Transformer, RINS): adds expressivity but does not solve state tracking — parallel propagation across positions still pushes state deeper.
  • Step-level recurrence is what truly enables unbounded state tracking, but it prevents parallelized training. Compromises include:
  • Mamba / SSMs: parallel training (convolution-like), RNN-like inference — but linear updates are no more expressive than a standard Transformer (Merrill et al., 2025).
  • DeltaNet: fast weight programming + delta rule; with eigenvalue ranges extended to negatives (Grazzi et al., 2025), it trains in parallel and exceeds standard Transformer expressivity.
  • RWKV-7: expressive dynamic state evolution.
  • COCONUT / Hierarchical Reasoning Model: latent, multi-cycle-per-token reasoning.
  • DeepMind's five research directions

    1. Hybrid models combining gated linear attention / gated delta nets with standard Transformer blocks. 2. Approximating state tracking in feedforward Transformers via training objectives and structural priors (a stopgap). 3. Coarse-grained recurrence at sentence or semantic-chunk level (e.g., Sentence Gestalt). 4. Representational alignment enabled by residual connections (e.g., Canon Layers), allowing variable-depth models to work without full retraining. 5. Multi-stage training: parallel feedforward pretraining, then recurrent fine-tuning, with truncated gradients and backpropagation through recurrence.

    Implications for agent developers

  • Don't worship long context: scaling from 128K to 1M tokens does not fix implicit state tracking.
  • CoT is not free: the thinking tax accumulates over long-horizon agents.
  • Architecture sets the ceiling: for long-range consistency, multi-hop reasoning, and dynamic environments, consider Mamba/RWKV-7/DeltaNet, external memory systems, or hybrid fast-feedforward + slow-recurrent designs.
  • The end goal is a "continuous inner monologue": a model that combines Transformer-style parallel training with RNN-style evolving internal state.
  • Reference papers cited

  • Mozer et al. (2026). *The Topological Trouble With Transformers.* Google DeepMind, arXiv:2604.17121.
  • Merrill & Sabharwal (2025). *The Expressive Power of Transformers with Chain of Thought.* ICLR.
  • Giannou et al. (2023). *Looped Transformers as Programmable Computers.* ICML.
  • Gu & Dao (2024). *Mamba: Linear-Time Sequence Modeling with Selective State Spaces.* ICML.
  • Schlag et al. (2021). *Linear Transformers Are Secretly Fast Weight Memory Systems.* NeurIPS.
  • Peng et al. (2025). *RWKV-7: Expressive Dynamic State Evolution.*
  • Hao et al. (2025). *COCONUT: Continuous Latent Thought.*
  • Allen-Zhu & Li (2025). *Canon Layers: Aligning Representations Across Steps.*

Tags

#transformers#deepmind#state-tracking#chain-of-thought#recurrent-neural-networks#mamba#rwkv#ai-architecture

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208060