You may have experienced this: ask an AI to read a 50-page contract and answer a question about a clause. It answers 'page 32'. You check page 32—the claim is wrong. Ask again, it gives a different answer. Ask a third time, it changes again.
This isn't the model being dumb. It's context rot—when the context gets too long, the model 'forgets' details it read earlier or mixes information from different parts. Hong et al. (2025) named this phenomenon.
Purbesh Mitra and Sennur Ulukus, in their paper *Chained Recursive Language Models for Multi-Iteration Reasoning*, propose a simple solution: don't make the model do everything in one shot. Let it work like colleagues handing off shifts—processing in multiple passes, each starting fresh, passing only the necessary 'work notes'.
The Problem: One Inference Does Too Much
Traditional long-context reasoning stuffs everything into the context and asks the model to simultaneously read, find evidence, track hypotheses, compute, use tools, and decide when to answer—like one person scanning a giant whiteboard left to right while remembering everything important.
Once the context grows long, the model starts:
- Shallow extraction: skimming surfaces without digging deep
- Forgetting what was checked: not remembering what was already verified
- Premature finalization: answering before truly auditing
- Root call 1: read pages 1–50, write AI-related statements to the blackboard, summarize: 'Processed 1–50, found 3 AI statements: 2 outlook, 1 promise.'
- Root call 2: receives summary + blackboard + artifacts, reads 51–100, adds 2 new statements, updates the blackboard.
- Root call 3: reads 101–150, flags a statement contradicting one from root call 1.
- Root call 4: reads 151–200, integrates everything, outputs the final answer.
- RULER: standard LLM $0.11/task, Chained RLM $0.21/task
- BABILong: standard LLM $0.14/task, Chained RLM $0.28/task
Chain-of-thought, self-consistency, and Tree-of-Thoughts all try to mitigate this, but operate within a single inference. Chained RLM operates at a different layer.
The Solution: Handoffs, Not Overtime
The core design of Chained RLM: the same model is invoked repeatedly, each invocation a fresh 'root call' that does not inherit the full previous conversation history.
Three handoff artifacts bridge iterations:
1. Summary: a compact plain-text description of 'what we've accomplished so far'—the briefing for the next shift. 2. Blackboard: a plain-text workspace recording key findings, open questions, and intermediate results. The next shift can read, modify, and extend it. 3. Artifacts: task-specific persistent products. For a codebase, a list of identified functions; for a long document, an extracted event timeline. More precise than plain-text summaries.
The raw predecessor trajectory is stored separately and, by default, not passed along—only fetched on demand, avoiding drowning the context in verbose history.
A Concrete Example
Task: given a 200-page earnings call transcript, find all management statements about AI investment and classify promises vs. outlook.
Traditional approach: stuff all 200 pages in, answer once—likely missing half or mixing quarters.
Chained RLM approach:
Each root call is 'fresh'—no fatigue from reading 150 pages, no attention dilution. Each restarts with lean work notes.
Results
Compared on four long-context benchmarks (both using GPT-5-mini):
| Benchmark | Standard LLM | Chained RLM | Gain | |-----------|-------------|-------------|------| | RULER | 87% | 92% | +5pp | | BABILong | 44% | 59% | +15pp | | LongBench v2 | 41% | 52% | +11pp | | OOLONG-real | 14% | 38% | +24pp |
A clear pattern: the more a task requires integrating distant information, the bigger the advantage. RULER gains are small because many tasks are solvable by local retrieval. OOLONG-real shows the largest gain because it demands retaining partial evidence, comparing distant events, and aggregating many small observations—exactly where single-pass inference fails.
The Cost: More Expensive, But More Auditable
Chained RLM isn't free: roughly 1.8x token consumption. Based on GPT-5-mini pricing:
Why It Matters
First, it shifts the long-context problem from 'stuff more' to 'split smarter'. The past two years focused on expanding windows from 32K to 1M, 10M. But even with infinite windows, single-pass inference rots. The question isn't 'can it fit' but 'can it be processed well once it fits'—a layer switch from scaling to divide-and-conquer.
Second, it echoes the 'granularity isomorphism' principle: the granularity of agent optimization should match the granularity of the object being optimized. Long-document processing naturally operates at the paragraph-level evidence-integration granularity, and Chained RLM aligns the reasoning unit to it.
Third, it makes intermediate reasoning auditable. Traditional one-shot inference is a black box. Every root call leaves a summary, blackboard, and artifacts—you can check how the model understood the first 100 pages and re-run from any checkpoint.
Honest Limitations
The researchers acknowledge: gains on RULER are limited—divide-and-conquer overhead isn't worth it for directly retrievable tasks. The doubled cost is real—not every scenario justifies paying twice for auditability and accuracy. And not all long-context benchmarks were tested; OOLONG-real's 14% → 38% is a big relative gain but the absolute value remains low.
But the value of this work isn't 'how much better than single-pass'—it points to a path different from 'keep expanding windows': make reasoning itself an engineering process that can be divided, handed off, and audited.
---
Paper: Chained Recursive Language Models for Multi-Iteration Reasoning, Mitra & Ulukus, 2026
Code: alexzhang13/rlm (GitHub)