A Counterintuitive Fact
You might think the bottleneck for coding agents like Claude Code and Codex is model capability. It's not. It's the context window burning money.
When an agent works autonomously in a large codebase—reading dozens of files, searching, editing, searching again—its context bloats every turn. Every additional token raises inference cost linearly, while model performance actually declines. This phenomenon has a name: context rot.
The Prime Intellect team put it bluntly in their January 2026 research blog: per-token cost grows linearly with context length, while model capability decreases with it. One curve rises, the other falls—their intersection is the death line of agent economics.
Prime Agent is their answer—an open-source coding agent that gained 2,271 GitHub stars in a single day, built around a core abstraction called Recursive Language Model (RLM).
RLM: Treat Context as a Variable, Sub-agents as Functions
RLM's core idea in one sentence:
> Don't have one LLM swallow all the context. Let it manage context as variables in a Python REPL and call sub-LLMs as functions.
Traditional agents work like this: all information goes into the prompt → LLM reasons → output. Context grows monotonically and can't shrink.
RLM works differently: the LLM gets a persistent IPython environment and can:
- Store large PDFs, datasets, and videos in Python variables instead of loading them into context, retrieving them with code when needed
- Call
rlm(...)to spawn sub-agents, delegating tasks; results return programmatically (not as natural-language summaries) - Sub-agents can spawn further sub-agents, forming a recursive call stack
- The base system prompt is immutable—
/refinenever rewrites the base prompt - Every update has a snapshot, supporting rollback
- Updates must be evidence-backed; no arbitrary changes
- The total token budget?
- Dependencies between sub-agents?
- Timeouts so one hung sub-agent doesn't freeze the whole tree?
- GitHub: https://github.com/PrimeIntellect-ai/prime-agent
- RLM blog: https://www.primeintellect.ai/blog/rlm
- RLM paper: https://arxiv.org/abs/2512.24601
- Continual Harness paper: https://arxiv.org/abs/2605.09998
This is not file-system-style external memory (like Claude Code's summarization). It's the model actively managing its own context window. The difference: external-file approaches are "use it and forget it"; RLM is "the model decides when to forget and when to call it back."
Why This Is More Than Engineering Optimization
The Prime Intellect blog makes a key claim:
> "teaching models to manage their own context end-to-end through reinforcement learning will be the next major breakthrough, enabling agents to solve long-horizon tasks spanning weeks to months."
The weight of this statement: they aren't building better scaffolding—they are training the model to learn scaffolding itself.
The core insight of the RLM paper (arXiv:2512.24601, Alex Zhang, October 2025): instead of writing ever more complex external context-management logic (summarization, compression, retrieval), let the model learn on its own, via RL, when to delegate, when to fold context, and when to call sub-agents.
This is in the same spirit as Sutton's "Bitter Lesson": hand-crafted scaffolding is the short-term fix; learning to scaffold yourself is the long-term one.
Continual Harness: Letting Agents Accumulate Experience
Prime Agent's second core abstraction is the Continual Harness. The problem it solves:
> Agents start every new session from zero and never accumulate experience.
The Continual Harness stores supplemental prompts, memories, skill descriptions, and subagent specifications as durable state. The agent can make small, incremental improvements to its own trajectory via a /refine command—evidence-backed updates, not arbitrary rewrites.
Key design constraints:
This suggests an analogy: the Continual Harness is git commits for an agent. The base prompt is the main branch, no direct pushes allowed; supplemental state is a feature branch, each refine is a commit—with diffs, history, and revert.
Why Prime Agent Gained 2,271 Stars in a Day
Three factors stacking up:
1. Novelty of the RLM paradigm: not better scaffolding, but training the model to scaffold itself. A paradigm-level bet. 2. Engineering completeness: daemon-backed sessions (no loss on disconnect), agent-to-agent communication (without routing through the user), heartbeat/schedule (timed wake-ups), autonomous mode (unattended operation). Not a demo—a production tool for long-running tasks. 3. Timing: in August 2026, the coding-agent market was converging on the consensus that "context rot is the #1 bottleneck." Prime Agent directly answers that consensus.
A Question Worth Thinking Hard About
RLM's recursive call stack means an agent can spawn sub-agents, which spawn grandchildren. Each level has its own context window. This is isomorphic to an operating system's process tree—init spawns services, services spawn workers, workers spawn ephemeral tasks.
But operating systems have schedulers, resource limits, and permission isolation. The RLM agent tree currently has none of these.
When agents can recursively spawn sub-agents without bound, who manages:
Prime Agent leaves these questions to the Continual Harness's /refine mechanism—letting the agent learn to manage them itself. But that "learning" requires massive RL training—and where does the training data come from?
This is the open question of the RLM paradigm, and the focus of Prime Intellect's next phase of research.
Links
One-sentence summary: Prime Agent isn't just a better coding agent—it's a bet that the next LLM breakthrough is "learning to manage its own context," not "a bigger context window."