English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

The Two-Hop Reasoning Paradox: LLMs Know Each Hop but Can't Combine Them

Forum topic · ✨步子哥 · 2026-08-10

Summary

A forum post on zhichai.net discusses a mechanistic interpretability paper (arXiv:2608.07261) explaining why large language models fail at two-hop reasoning even when they can answer each hop correctly. The paper's authors trained transformers from scratch on controlled symbolic knowledge-graph data and found that when the second-hop relation is in-distribution, accuracy is nearly 100%, but it drops to about 30% out-of-distribution. Probing experiments show the bridging entity and even the correct final answer exist in intermediate-layer representations, yet upper layers perform a memorized 'bridging entity → answer' lookup (mapping) rather than relational inference, so they fail on unseen bridge entities. The authors propose two fixes: training upper layers explicitly to reason over representations, and a Looped Transformer architecture with shared parameters that re-processes intermediate representations across iterations, raising OOD two-hop accuracy from roughly 30% to 80%+ without hurting in-distribution performance. The post argues this diagnosis may underlie Chain-of-Thought effectiveness and RAG failures where correct evidence is retrieved but not used, while noting limitations: the symbolic setup, unclear benefits of parameter sharing, and the gap between probe-detectable and actually-used information.

preview _4.svg

> Paper link: https://arxiv.org/abs/2608.07261

A Puzzling Failure

Ask an LLM: "Which university did the wife of the author of *One Hundred Years of Solitude* attend?"

This is two-hop reasoning:

  • Hop 1: Who is García Márquez's wife? (Mercedes Barcha)
  • Hop 2: Which university did Mercedes Barcha attend?
  • The model can answer each hop correctly in isolation — but ask the combined question and it starts hallucinating.

    The model knows each hop. Why can't it combine them?

    This is the "two-hop generalization paradox" the paper sets out to explain. Rather than settling for "models are bad at reasoning," the authors dig into the transformer's internal mechanisms and find a rather counterintuitive answer.

    Experiments: Starting from a Controlled Setting

    To rule out pretraining-data interference, the authors trained transformers from scratch in a fully controlled symbolic knowledge-graph environment. Each training sample is a triple (head entity, relation, tail entity), and the model must learn two-hop queries: given a head entity and two relations, find the final tail entity.

    Key finding: when the second hop's relation pattern is in-distribution (ID), the model is nearly 100% correct; when the second hop is slightly out-of-distribution (OOD), accuracy drops to about 30%.

    Note: hop 1 is always correct — the model finds the right bridging entity. The problem is the second hop: even with the correct bridge entity present in context, the model fails to use it.

    Mechanism: Upper Layers Do "Mapping," Not "Reasoning"

    This is the most interesting part of the paper. Using mechanistic interpretability tools:

    Step 1: Locating the bridging entity's representation. Logit lens and patching experiments show the bridge entity is correctly represented in middle layers (roughly layers 4–8). The model "sees" it; it exists in the hidden space.

    Step 2: Probing. A linear probe trained on intermediate representations can predict the correct tail entity. The correct answer's information is present in the intermediate layers — it's just sitting there.

    Step 3: Why doesn't it come out? The key finding: the upper (final) layers perform "mapping," not "reasoning."

    In-distribution, the upper layers learn a direct "bridge entity → tail entity" lookup table. This mapping is static — not derived from relational reasoning, but hardcoded from (bridge, answer) pairs seen during training. So once an unseen bridge entity appears, the lookup fails.

    Analogy: you memorize 100 two-digit addition problems, including 23+45=68 and 12+34=46. On the exam, 23+45 is fine (lookup hit) but 23+46 stumps you (lookup miss) — even though you "know" 3+6=9 and 20+40=60. You memorized answers, not an algorithm.

    The transformer's upper layers are that student. In-distribution they memorize a lookup table; faced with new bridge entities, they lack the ability to reason out the answer.

    Solutions: Making the Upper Layers Actually Reason

    Option 1: Explicitly train upper layers for representation-based reasoning. Freeze lower layers and train only the upper layers to infer the tail entity from the bridge entity's intermediate representation rather than look it up. It works but requires carefully designed training data.

    Option 2: Looped Transformer. Share parameters across the upper layers, letting the model "think it over" multiple times. Each loop re-processes the intermediate representations, forcing the upper layers to learn a general reasoning procedure rather than a fixed mapping.

    Results: the Looped Transformer lifts OOD two-hop accuracy from ~30% to 80%+ with no loss on ID performance. Logit lens analysis confirms the looped upper layers perform representation-based reasoning, not a bigger lookup table.

    Why This Paper Matters

    1. It pinpoints a concrete mechanistic defect. Not the vague "poor reasoning ability," but an actionable finding: upper layers do mapping, not reasoning.

    2. The Looped Architecture has deep implications. It essentially makes the transformer emulate repeated processing — like iterative deliberation — re-extracting features each round. It resembles an RNN's unrolling over time, but motivated differently: not for sequences, but to give reasoning "thinking time."

    3. It informs how we understand Chain-of-Thought. CoT may work not because the model "verbalizes intermediate steps," but because generating them forces re-representation in context, giving upper layers a second reasoning pass. The Looped Architecture internalizes this — multi-round reasoning without external CoT.

    4. Direct relevance to knowledge-graph-augmented LLMs. RAG systems often "retrieve the right evidence but the model can't use it." This paper's diagnosis — upper layers map instead of reason — may be the underlying cause of such failures.

    An Honest Assessment

    Limitations:

  • The experimental environment is symbolic and controlled. Real-world two-hop reasoning involves natural language, entity disambiguation, and relation ambiguity. The appendix validates on real datasets, but at limited scale.
  • The benefit mechanism of the Looped Architecture needs deeper explanation. Is parameter sharing acting as regularization, or something deeper? Evidence is suggestive but not conclusive.
  • The "mapping vs. reasoning" distinction rests on probing. Information findable by a probe ≠ information the model actually uses. More rigorous causal intervention experiments are needed.
Still, this is elegant mechanistic interpretability work — not just "opening the black box for a look," but "locating a specific computational defect and fixing it." This diagnose-then-repair loop is what mechanistic interpretability should aspire to.

---

Paper: Zhang, Wang, Wang, Wan, Luo. *Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models*. arXiv:2608.07261, 2026.

Tags

#llm-reasoning#mechanistic-interpretability#transformers#looped-transformer#chain-of-thought#two-hop-generalization#knowledge-graphs#rag

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633311