A Lab That Never Sleeps
Imagine a lab with four researchers: a Manager deciding what the campaign does next, a Planner breaking goals into bounded tasks, an Engineer running experiments and writing code, and a Reviewer independently checking results before anything is declared done.
This lab has no breaks. It can work on a single math problem for days, survive seven wrong turns without collapsing, and record every rejected path so future rounds don't repeat it. Six papers from proposal to submission, 254 subtasks, 16 rollbacks — all self-directed.
This isn't a bigger language model. It's a runtime, and its name is Argus.
The Core Problem: Long-Horizon Reasoning Isn't "Longer Reasoning"
An agent can call tools 200 times in a row, but if it never retains lessons from the first 100 failures, call 101 is identical to call 1. The real bottleneck isn't context window length — it's state management:
- When to persist? The current method seems to work, but maybe you haven't hit the wall yet.
- When to pivot? Failure information may hide the correct direction.
- When to admit defeat? Better to record "this path is blocked" than fabricate a completion.
- Control plane: anchors the campaign, dispatches tasks
- Execution plane: runs one bounded task
- Record plane: stores what happened, but has no completion authority
- Ramp-up (W1–6): minimal project-specific state
- Mature (W19–22): reusable knowledge accumulated
- Tail (W23–24): hardest tasks
- Forward latency: 0.199ms → 0.168ms (1.18x)
- Forward+backward: 0.900ms → 0.747ms (1.21x)
- 13/13 correctness gates passed; 14/14 repo tests passed
- Division beats unification: four specialized roles echo the "small model screens, big model verifies" pattern.
- Granularity isomorphism: campaign → mission → round hierarchy matches the insight that optimization granularity should match the behavior being optimized (trajectories, not single calls).
- Evaluation blind spots: "once correct ≠ currently correct" is exactly what the Reviewer revision funnel addresses.
- Solving at a different layer: the model generates; the runtime coordinates, reviews, and accumulates — like letting the LLM be the poet and a solver be the accountant.
- Models are swappable (GPT-5.5 → GLM-5.2, runtime keeps running)
- Sessions can break (restart recovers from checkpoint)
- State is auditable (it's plain text, readable by humans and Reviewers)
- Ramp-up vs. mature comparison is observational, not causal (no frozen-state matched replay)
- Reviewer routing is adaptive, not random — no direct causal attribution
- The six papers come from one environment with overlapping campaign-hours
- GLM-5.2 on Claude Code is in progress: 70.94% without a matched Direct baseline
- Chip results stop at synthesis and static timing — no place-and-route, power, or tape-out
- The MOF 989-crystal subset uses the same data as design scoring, not an independent holdout
- arXiv: https://arxiv.org/abs/2608.05144
- Full HTML: https://arxiv.org/html/2608.05144v1
- Upstream merge: FLA PR #1045
All three reduce to one need: a memory layer more durable than model weights — but one that can't become an unreviewed garbage heap.
Argus's Answer: Four Roles + Three Planes + Verification Gating
Argus's thesis in one line: model weights stay fixed; runtime state evolves.
Four Asymmetric Roles
| Role | Responsibility | Authority | |------|----------------|-----------| | Manager | Campaign-level goal management | Only role that can change phases | | Planner | Converts state into the next bounded task | Owns task definition | | Engineer | Executes tasks, produces artifacts | May self-review low-risk tasks | | Reviewer | Independent checks, revision, or veto | Final gate on high-risk tasks |
The novelty isn't the division of labor — it's the asymmetric authority: Engineer can self-approve small tasks, but any phase closure, vertical strategy, or help request triggers mandatory Reviewer involvement. Crucially, the Reviewer reads the Engineer's actual artifacts, not the Engineer's summary of them — like a construction inspector measuring wall thickness instead of reading the build log.
Three-Plane Separation
Denying the record plane completion power prevents collusion between the recorder and the executor. Recording is just recording — it never becomes "done" because the Engineer says so.
Verification Gating
A candidate update — memory, skill, validation rule, or rejected path — is never adopted merely because a role produced it. It must pass:
1. Task-native evidence checks (tests, verifiers, formal checks) 2. Submission by an authorized role (Engineer self-review or Reviewer) 3. Retrievable reuse in subsequent tasks
Only the full submit-then-reuse path counts as runtime self-evolution.
Benchmark Results: Seven Arenas
The authors explicitly refuse to average scores across benchmarks with different units.
| Benchmark | Argus | Reference | |------|-----------|---------| | SWE-Bench Pro | ~78% | 59% (Direct Copilot) | | AARRI-Bench | 76.8% | — | | Math data synthesis | 28.0-point gap | — | | GPU kernel optimization | Competitive | — | | LM training | Competitive | — | | Research assistant | Competitive | — | | Data synthesis | Competitive | — |
The SWE-Bench Pro gap is +19 points at 1.41x token cost — the extra tokens go to planning, execution, and review, which is exactly what buys the improvement.
Evidence of Runtime Self-Evolution
Splitting SWE-Bench Pro's 731 tasks into execution-order windows:
Mature windows use 21% fewer solve tokens and 15% less active time per task than ramp-up — the same model gets more efficient because the runtime remembers which paths not to walk. The authors are honest: the curve isn't monotonic (W13–18 had lowest tokens but longer latency), and hard tasks don't get easier just because the runtime matured.
The Reviewer Repair Funnel
Of 731 tasks: 466 (63.7%) routed to independent Reviewer, 265 (36.3%) self-reviewed. The Reviewer demanded revisions on 43 tasks; 34 passed official verifiers after revision; 22 completed full revise-to-done loops. Reviewer-routed tasks consumed 2.75x tokens and 1.80x time — routing concentrates review on harder tasks, not randomly. The Reviewer is a correction channel, not a post-hoc commentator: 79.1% of revised solutions passed official verification.
RWKV6 Kernel: A Real, Human-Reviewed Upstream Contribution
Argus optimized a TileLang RWKV6 kernel and submitted it to the Flash Linear Attention repo (fla-org). A Moonshot AI collaborator reviewed the generated CUDA code, found a long-sequence numerical stability issue, and required block-local exponent centering. Argus fixed it and re-ran the validation suite; it was merged to main on July 20, 2026.
This is human-reviewed upstream adoption, not self-reported success.
Math Campaign: Keeping a Rejected Path
In a multi-day Erdős–Gyárfás number campaign: the Manager recovered the campaign after a budget failure; the Planner first chose a cheap falsification test, then changed the success criterion from "collect observations" to "produce a proof"; the Engineer searched literature and wrote executable checks; the Reviewer vetoed an overclaiming path, demanded extra checks, and accepted a tighter bound.
The final state preserved one rejected path and six proved frontier updates. A rejected path isn't failure — it's reusable knowledge that lets later tasks skip it, like a well-kept lab notebook of failed experiments.
Six Papers: 254 Tasks, 16 Rollbacks
Argus ran six full research projects (evaluation reliability, vision-language matching, test-time adaptation, GUI agents, multimodal hallucination, model quantization), all reaching submission. The point isn't that six papers got written — it's that not one went through in a single pass: 286 Reviewer revision verdicts, 89 session switches, 16 stage rollbacks.
The multimodal hallucination project is most telling. In a 12-hour search window, seven rollback decisions rejected method paths for incomplete baseline coverage, exact duplication of a baseline, or missing pre-registered signals. The Planner then reframed "positive mitigation" as a diagnostic negative-result study; the Manager accepted the redefinition; the Engineer delivered a 5-method × 3-benchmark matrix with 4,500 lines of official scoring data; the Reviewer bound "ineffective" and "degenerate" conclusions to those outputs.
Failure is preserved as information; only reviewed state drives the next step.
Engineering Insight: Why "Fixed Weights" Wins
Argus's counterintuitive conclusion: you may not need a bigger model for better long-horizon reasoning — you need a better runtime.
CHECKPOINT.md: The Humblest, Most Critical Design
Every Engineer/Reviewer call uses a fresh provider session. Cross-session continuity comes from a plain CHECKPOINT.md file holding persistent state, evidence references, open questions, and next steps.
This moves memory out of model context and into the filesystem:
Don't let the context window become an ever-lengthening tail; make it a projection that can be rebuilt at any time.
Honest Limitations
The paper's limitations section deserves note:
This honesty itself embodies verification gating: never claim what the verifier didn't cover.
Personal Take: The Runtime Is the New "Operating System"
Argus strengthens a conviction: the "operating system" of the agent era isn't the model — it's the runtime.
The model is the CPU: it executes but doesn't schedule. The runtime is the OS: processes, memory, filesystem, permissions. The three-plane split maps to scheduling/execution/logging; four roles map to process/thread/file/permission managers; verification gating maps to permission checks; CHECKPOINT.md maps to persistent storage.
The analogy is actionable: agent competitiveness lies in runtime design quality, not parameter count. An 8B model with a good runtime may beat a 70B model with a bad one on long-horizon tasks — because the latter keeps repeating its mistakes and the former doesn't.
Lightweight, persistent, auditable runtimes are the next main battleground of agent engineering.
Paper Links
*Argus is the hundred-eyed giant of Greek myth who never fully sleeps. The name is precise: a runtime with part of its eyes always open, watching its own every step.*