> Source: zhichai.net forum post, an in-depth Chinese-language analysis of the MIT CSAIL blog post "Language model harnesses are compositional generalizers" by Alex Zhang and Omar Khattab (published July 20, 2026, as a blog post — not peer-reviewed). > Original: https://alexzhang13.github.io/blog/2026/harness > Lineage: Recursive Language Models (2025-10 blog / arXiv:2512.24601) → Mismanaged Geniuses Hypothesis (2026-04) → this post (2026-07).
Key points
- Title correction: The actual title is "Language model harnesses are compositional generalizers" — the subject is the *harness*, not the model. The core claim: what generalizes is not the neural network but the program wrapping it.
- Thesis: Don't fix the network, fix the harness. The harness (formalized H: s → a) decides what the model sees, and can carry symbolic, structural inductive biases that token-level Transformer biases lack.
- Compositional generalization = recombining learned parts to solve unseen problems. Fodor & Pylyshyn (1988) called this gap "systematicity"; it remains unsolved.
- The hardest evidence comes from *Faith and Fate: Limits of Transformers on Compositionality* (Dziri et al., NeurIPS 2023):
- Transformers solve compositional tasks via *linearized subgraph matching* — retrieving training paths, not executing algorithms. Error rates compound exponentially over autoregressive steps.
- The MIT authors: generalization "has been left to the underlying neural network and to its 2017 token-level inductive biases," and scaling returns on compositional ability are "just not budging."
- Practical consequence: length generalization failure — training at one context length doesn't transfer to longer ones, forcing every model generation to re-pay the long-context training tax.
- Metaphor: don't make a Jiangnan-cuisine chef cook an entire imperial banquet; decompose it into a hundred steps he already knows, and orchestrate them with a workflow.
- Mainstream agents violate LID: ReAct, CodeAct, Claude Code, Codex all append observations to a growing prefix (
context += action + obs), so by loop 50 the prompt looks like nothing in the training data. This reframes *context rot*: the poison is not length itself but the distributional unfamiliarity that length produces. - Equivalence classes: a harness induces an equivalence relation over task states, so structurally similar tasks render token-for-token similar contexts — generalization is constructed, not learned.
- Recursive Language Models keep a giant context in a REPL (e.g., 5,000,000 characters visible to the root model as ~20 tokens of metadata) and recursively call the same model on decomposed slices.
- Context offloading: store content in variables, not prompts; programmatic sub-calls: code (with for-loops, variables, recursion) spans a richer space of decompositions than discrete tool calls.
- Reported result: training on short tasks transfers to million-token and cross-domain tasks because the harness quotients trajectories onto a shared in-distribution structure.
- Design agents so each LLM call sees a familiar-shaped input; avoid unbounded context appending.
- Treat the harness as a first-class carrier of inductive bias, not glue code.
- Consider REPL-style variable offloading and recursive programmatic sub-calls for long-context workloads.
Why Transformers fail at compositional generalization
| Experiment | Result | |---|---| | GPT-4, 3-digit × 3-digit multiplication | 59% | | GPT-4, 4-digit × 4-digit | 4% | | 3×3 with few-shot scratchpad | 59% → 92% | | Scratchpad on high-complexity problems | still near 0% | | GPT-3 fine-tuned to 100% on 1.8M 3×3 examples, tested on 4×4 | 0% | | 4×2 cases where answer is right but intermediate steps wrong | 82% ("restoration error" — recall, not computation) |
The LID criterion (Locally In-Distribution)
The article's most valued idea:
> A good harness produces observations that are *locally in-distribution*: every individual LM call is in-distribution with respect to the training data — even if the overall task is OOD.
RLMs: context as environment, not input
Caveats raised by the author
1. It's a blog post, not a peer-reviewed paper — read it as frontline experimental notes, not settled science. 2. The "isn't this just subagents?" objection has real substance. 3. Generalization isn't automatic; it needs the right harness design push. 4. Real costs: more LM calls mean higher latency and token expense. 5. Theoretical concern: risk of circular reasoning if the harness's notion of "in-distribution" is defined by what it itself produces.