> Paper link: https://arxiv.org/abs/2608.07261
A Puzzling Failure
Ask an LLM: "Which university did the wife of the author of *One Hundred Years of Solitude* attend?"
This is two-hop reasoning:
- Hop 1: Who is García Márquez's wife? (Mercedes Barcha)
- Hop 2: Which university did Mercedes Barcha attend?
- The experimental environment is symbolic and controlled. Real-world two-hop reasoning involves natural language, entity disambiguation, and relation ambiguity. The appendix validates on real datasets, but at limited scale.
- The benefit mechanism of the Looped Architecture needs deeper explanation. Is parameter sharing acting as regularization, or something deeper? Evidence is suggestive but not conclusive.
- The "mapping vs. reasoning" distinction rests on probing. Information findable by a probe ≠ information the model actually uses. More rigorous causal intervention experiments are needed.
The model can answer each hop correctly in isolation — but ask the combined question and it starts hallucinating.
The model knows each hop. Why can't it combine them?
This is the "two-hop generalization paradox" the paper sets out to explain. Rather than settling for "models are bad at reasoning," the authors dig into the transformer's internal mechanisms and find a rather counterintuitive answer.
Experiments: Starting from a Controlled Setting
To rule out pretraining-data interference, the authors trained transformers from scratch in a fully controlled symbolic knowledge-graph environment. Each training sample is a triple (head entity, relation, tail entity), and the model must learn two-hop queries: given a head entity and two relations, find the final tail entity.
Key finding: when the second hop's relation pattern is in-distribution (ID), the model is nearly 100% correct; when the second hop is slightly out-of-distribution (OOD), accuracy drops to about 30%.
Note: hop 1 is always correct — the model finds the right bridging entity. The problem is the second hop: even with the correct bridge entity present in context, the model fails to use it.
Mechanism: Upper Layers Do "Mapping," Not "Reasoning"
This is the most interesting part of the paper. Using mechanistic interpretability tools:
Step 1: Locating the bridging entity's representation. Logit lens and patching experiments show the bridge entity is correctly represented in middle layers (roughly layers 4–8). The model "sees" it; it exists in the hidden space.
Step 2: Probing. A linear probe trained on intermediate representations can predict the correct tail entity. The correct answer's information is present in the intermediate layers — it's just sitting there.
Step 3: Why doesn't it come out? The key finding: the upper (final) layers perform "mapping," not "reasoning."
In-distribution, the upper layers learn a direct "bridge entity → tail entity" lookup table. This mapping is static — not derived from relational reasoning, but hardcoded from (bridge, answer) pairs seen during training. So once an unseen bridge entity appears, the lookup fails.
Analogy: you memorize 100 two-digit addition problems, including 23+45=68 and 12+34=46. On the exam, 23+45 is fine (lookup hit) but 23+46 stumps you (lookup miss) — even though you "know" 3+6=9 and 20+40=60. You memorized answers, not an algorithm.
The transformer's upper layers are that student. In-distribution they memorize a lookup table; faced with new bridge entities, they lack the ability to reason out the answer.
Solutions: Making the Upper Layers Actually Reason
Option 1: Explicitly train upper layers for representation-based reasoning. Freeze lower layers and train only the upper layers to infer the tail entity from the bridge entity's intermediate representation rather than look it up. It works but requires carefully designed training data.
Option 2: Looped Transformer. Share parameters across the upper layers, letting the model "think it over" multiple times. Each loop re-processes the intermediate representations, forcing the upper layers to learn a general reasoning procedure rather than a fixed mapping.
Results: the Looped Transformer lifts OOD two-hop accuracy from ~30% to 80%+ with no loss on ID performance. Logit lens analysis confirms the looped upper layers perform representation-based reasoning, not a bigger lookup table.
Why This Paper Matters
1. It pinpoints a concrete mechanistic defect. Not the vague "poor reasoning ability," but an actionable finding: upper layers do mapping, not reasoning.
2. The Looped Architecture has deep implications. It essentially makes the transformer emulate repeated processing — like iterative deliberation — re-extracting features each round. It resembles an RNN's unrolling over time, but motivated differently: not for sequences, but to give reasoning "thinking time."
3. It informs how we understand Chain-of-Thought. CoT may work not because the model "verbalizes intermediate steps," but because generating them forces re-representation in context, giving upper layers a second reasoning pass. The Looped Architecture internalizes this — multi-round reasoning without external CoT.
4. Direct relevance to knowledge-graph-augmented LLMs. RAG systems often "retrieve the right evidence but the model can't use it." This paper's diagnosis — upper layers map instead of reason — may be the underlying cause of such failures.
An Honest Assessment
Limitations:
---
Paper: Zhang, Wang, Wang, Wan, Luo. *Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models*. arXiv:2608.07261, 2026.