English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MACRO: Rerouting Transformer Layers with Markov Chains Boosts Accuracy by 26 Points Without Touching Weights

Forum topic · ✨步子哥 · 2026-08-08

Summary

MACRO, a paper by Batorskq et al. (August 2026, arXiv:2608.05872), proposes rerouting the layer execution order of Transformers via Markov chain modeling instead of modifying model weights. Building on Dr. LLM—which allowed skipping or repeating layers but required 14.8 hours of brute-force search per route—MACRO models layer-jump decisions as a Markov chain and uses Viterbi dynamic programming to find optimal execution paths, cutting search time to 1.6 hours (9.4x faster). Tested across 5 models (Llama3-8B, Qwen3-1.7B/4B/7B, Mistral-7B) and 7 benchmarks (GSM8K, MATH500, OpenBookQA, ARC-C/E, CommonsenseQA, HellaSwag), MACRO achieves +5.0% average improvement over baselines and +7.2% over Dr. LLM. Notably, Qwen3-1.7B on GSM8K rises from 43.4% to 69.5% by rewinding layer 7 to layer 3. Routes transfer across same-type tasks, and a single shared route trained on multiple tasks generalizes to unseen tasks (+9.02% on out-of-domain benchmarks). Mechanistic analysis shows correct-answer signals exist in intermediate layers but are suppressed toward output; routes preserve them. Gains diminish with model size, suggesting MACRO is most valuable for small models.

Have you ever changed a correct exam answer back to a wrong one before submitting? You knew the answer—the reasoning step just led you astray. Large language models have the same flaw: they often internally "know" the correct answer, but the standard left-to-right, layer-1-to-layer-N forward pass suppresses that signal in certain layers.

In August 2026, Batorskq et al. published MACRO, a paper with a bold idea: without touching model weights, simply changing the layer execution order can rescue the suppressed correct signal.

Paper: https://arxiv.org/abs/2608.05872 Code: https://github.com/Batorskq/MACRO

1. The Problem: The Model Knows, the Forward Pass Derails It

A curious phenomenon has been observed in recent years: in the middle layers of a Transformer, the model's prediction probability for the correct answer is sometimes *higher* than at the final layer. The information exists in the middle layers, but subsequent processing "washes it out."

A prior method, Dr. LLM, tried to address this by allowing certain layers to be skipped or repeated at inference time, searching by trial and error for a better layer execution path (called a *route*). It worked well, but had a fatal drawback: the search space was so large that finding one route took 14.8 hours.

MACRO's entry point: model the layer execution path as a Markov chain. This is not just new mathematical clothing—it fundamentally changes the search method.

2. Core Idea: Layer Execution Path = Markov Chain

Standard Forward Pass vs. MACRO

A standard Transformer executes layers in a fixed order: layer 1 → layer 2 → layer 3 → ... → layer N. Each layer's output feeds only the next layer.

MACRO allows detours: layer 1 → layer 2 → layer 3 → layer 3 (repeat) → layer 4 → ..., or layer 1 → layer 2 → layer 7 → layer 8 → ... (skipping intermediate layers).

The key innovation is how the path is searched. Dr. LLM uses brute force: trying every combination of "skip / repeat / normal" per layer—an exponential number of combinations. MACRO converts this into probabilistic optimization over a Markov chain.

How the Markov Chain Helps

The core Markov assumption: the next state depends only on the current state, not earlier history. MACRO models each layer's jump decision (which layer to jump to) as a probability distribution, then uses the Viterbi algorithm (dynamic programming) to find the optimal path instead of brute-force enumeration.

The benefit is enormous: search time drops from 14.8 hours to 1.6 hours—a 9.4x speedup. More importantly, because the search space is structured, the routes found are of higher quality.

3. Results: Numbers Speak

Main Experiments

The paper evaluates 5 models (Llama3-8B, Qwen3-1.7B/4B/7B, Mistral-7B) on 7 benchmarks (GSM8K, MATH500, OpenBookQA, ARC-C, ARC-E, CommonsenseQA, HellaSwag).

Key numbers:

  • +5.0% average improvement over the non-routed baseline
  • +7.2% over Dr. LLM on average
  • 9.4x faster search (14.8h → 1.6h)
The most striking single result: Qwen3-1.7B on GSM8K jumps from 43.4% to 69.5%—a 26-point gain just by "rewinding" from layer 7 back to layer 3. No fine-tuning, no extra data, no architecture change—only a different layer execution order.

Route Transfer: Math Routes Transfer to Math

MACRO finds an interesting pattern: a good route found on GSM8K (math word problems) also improves MATH500 (competition math). But a route found on CommonsenseQA does not transfer to math reasoning.

This means routes are not random noise—they capture structural features of task types. Math reasoning requires a specific inter-layer information flow pattern, shared across different math tasks.

One Model, One Path

The most practical finding: a shared route trained jointly on multiple tasks transfers to completely unseen tasks. Qwen3-1.7B gains +9.02% on out-of-domain benchmarks on average. You don't need to search a new route per task—one general route covers a broad class of tasks.

4. Why It Works: Mechanistic Interpretability

The paper's deepest section is its mechanistic analysis. The authors find a key phenomenon:

In the standard forward pass, the correct-answer signal exists in middle layers but is increasingly suppressed toward the output. MACRO's routes preserve the correct signal to the end by skipping "suppressing" layers or repeating "amplifying" layers.

This is consistent with earlier findings on Othello-GPT and Toy Models of Superposition (TMS): Transformer middle layers encode rich information, but the standard forward pass is not necessarily the best way to use it.

The route effect also depends on model size: the smaller the model, the larger the gain. Llama3-8B benefits less than Qwen3-1.7B. This suggests large models may have already internalized route-like functions in their weights, while small models need routes to "externalize" them.

5. Conceptual Positioning: Solving Problems at a Different Level

MACRO joins a lineage of "solve it at a different level" concepts: octopus RNA editing (edit RNA, not DNA), slime mold externalized memory, avian quantum magnetoreception, SOPHIA's division of labor, EvoThink's atomic reasoning, Möbius RoPE topological intervention, mantis shrimp phononic shields, Euclid-MCP reasoning outsourcing, the Regression Tax paired evaluation, and ACE context engineering.

MACRO is the eleventh: don't change the weights, change the execution path. It is isomorphic to octopus RNA editing—the octopus doesn't change the blueprint (DNA), only the construction plan (RNA editing sites); MACRO doesn't change the blueprint (weights), only the construction plan (layer execution order).

6. Limitations and Honest Assessment

1. Route search still takes 1.6 hours: much faster than Dr. LLM, but too slow for real-time use. However, once found, inference is as fast as a standard forward pass. 2. Transfer has boundaries: math routes transfer to math tasks, but cross-type transfer (math → commonsense) fails. Route generality has limits. 3. Diminishing returns at scale: larger models gain less. MACRO may be most valuable for small models—precisely those most needed in resource-constrained settings. 4. Viterbi approximation: the Markov assumption (next layer depends only on current layer) is approximate, not exact. Experiments show it works in practice, but theoretically some non-Markov path patterns could be missed.

7. The Bigger Picture

MACRO points to a more fundamental question: is the standard Transformer forward pass optimal?

For years, nearly all methods for improving Transformers changed the weights: pretraining, fine-tuning, RLHF, DPO. MACRO offers another path: keep the weights, change the execution route—like getting a major performance boost from the same hardware by only changing the software scheduling policy.

This also relates to the "evaluation blind spot" idea: we always assumed the standard forward pass is the best way to use a Transformer, but never verified the assumption. MACRO shows it's wrong—at least for some tasks and models, better execution paths exist.

More deeply, MACRO hints that Transformer middle layers may be "doing" far more than we ask them to "say." The standard forward pass forces all information to be compressed into the final layer's output, but the middle layers may encode richer, non-linear reasoning paths. MACRO's routes are one way to unlock those capabilities.

8. Conclusion

MACRO's core insight: the model already knows the answer—the standard forward pass just leads it astray. Finding the right layer execution path rescues the suppressed signal.

It recalls the old joke about a man searching for lost keys under a streetlight—not because the keys are there, but because the light is good. The standard Transformer forward pass is that streetlight: we use it because it's convenient, not because it's optimal. MACRO reminds us the keys may be elsewhere.

Once we stop being bound by the assumption that layers must execute in order from 1 to N, Transformers may reveal capabilities we've never seen. MACRO opens one door; behind it may be an entire corridor.

---

Paper: https://arxiv.org/abs/2608.05872 Code: https://github.com/Batorskq/MACRO HTML version: https://arxiv.org/html/2608.05872v1

Tags

#llm#transformers#markov-chains#inference-optimization#mechanistic-interpretability#layer-routing#small-language-models#macor

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178603070