The Truth About Million-Token Context Windows: How RLMs Solve AI's Long-Context 'Dementia'
Key points
- Long context ≠ strong reasoning. Despite million-token windows, frontier LLMs degrade sharply on long-document reasoning tasks (e.g., cross-chapter financial report analysis), often acting like 'repeaters' that only restate surface information.
- Context Rot. MIT researchers show model performance drops significantly as input length and task complexity increase—even for GPT-5-class models. Causes include attention dilution (signals drowned in noise over long sequences), positional encoding limits, and a 'phase transition': a sudden collapse from decent memory-level performance to near-random output once complexity crosses a threshold.
- Recursive Language Models (RLM), proposed by MIT CSAIL, flip the paradigm: the long context is stored *outside* the model as a variable in a persistent Python REPL environment. The model becomes a manager/agent rather than a memorizer.
- OOLONG tests reasoning and aggregation over long text, not just needle-in-a-haystack retrieval. OOLONG-Pairs requires pairwise reasoning over all input entries—quadratic, O(N²), complexity.
- GPT-5 directly on OOLONG-Pairs: F1 ≈ 0.04% (effectively random).
- RLM (GPT-5-based): F1 = 58.00%; an ablated RLM *without* recursive calls still scored 43.93%, showing both context externalization (~large gain) and recursion (+14 points) matter.
- Cost: RLM selectively processes only relevant snippets, so it was often the *cheapest* method (e.g., $0.11 on CodeQA vs. $1.31 for a summary agent; $0.99 on BrowseComp+ vs. an estimated $1.50–$2.75 for direct GPT-5).
- Financial analysis: strategic section location, per-section recursive analysis, cross-validation and trend/risk synthesis—turning a 'repeater' into an analyst.
- Codebases: architecture mapping, dependency analysis, code review, vulnerability detection on large repositories.
- Multi-document summarization, legal research, scientific literature review: recursion supports cross-document aggregation and synthesis.
- Neuro-symbolic framing: the LLM supplies intuition/semantic understanding; the Python REPL supplies deterministic logic and precise control. RLM also makes reasoning more interpretable (visible, step-by-step code and calls), representing a shift from memorization toward structured thinking—a candidate direction toward more general AI.
How RLM works
1. Load context into a REPL: the full text is loaded as a variable (e.g., context) with metadata hints (length, type) given in the prompt.
2. Code-driven filtering: the model writes Python to peek (context[:1000]), search (regex), partition, and store relevant snippets—never ingesting everything into its own window.
3. Recursive sub-model calls: a special function like llm_query(prompt, sub_context) delegates small subtasks (summarize a chapter, verify a pair) to sub-models, which can themselves recurse, forming an adaptive task-decomposition tree.
4. Aggregation: the root model collects and integrates sub-results, emitting a final answer via markers like FINAL(answer) or FINAL_VAR(name).
Benchmark results (OOLONG / OOLONG-Pairs)
| Method | CodeQA (23K–4.2M tok) | BrowseComp+ (1K) (6M–11M tok) | OOLONG (131K tok) | OOLONG-Pairs (32K tok) | | :--- | :--- | :--- | :--- | :--- | | Base Model | 20.00%* | 0.00%* | 44.00% (GPT-5) | <0.1% (GPT-5) | | Summary Agent | 58.00% ($1.31) | 70.47% ($0.57) | 46.00% ($0.13) | 0.01% ($0.13) | | RLM (no recursion) | 58.00% ($0.18) | 88.00% ($0.44) | 36.00% ($0.37) | 43.93% ($0.69) | | Full RLM | 62.00% ($0.11) | 91.33% ($0.99) | 56.50% ($0.43) | 58.00% ($0.33) |
*Inability to fit input; costs are average API cost in USD (per the paper's Table 1, as cited in the post).*