Paper: Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks arXiv: 2608.05144 Authors: Boxiu Li, Zimo Wen, Yijia Fan, et al. (Shanghai Jiao Tong University, Microsoft, Fudan, Tsinghua, etc.)
Reference: https://arxiv.org/abs/2608.05144
The Core Problem
When today's agents tackle long-horizon tasks, the goal is often unclear from the start. As they run, they either "play dumb" (goal drift) or get stuck in loops. Argus targets scenarios where the goal itself must co-evolve with the problem being solved — not exam questions, but genuine research.
The Four-Role Architecture
| Role | Responsibility | |------|----------------| | Manager | Anchors user intent; owns authority over working-contract changes | | Planner | Decomposes the next task unit | | Engineer | Executes, implements, evaluates | | Reviewer | Audits artifacts and issues completion verdicts |
Key Design — The Working Contract Kt = (ι, ot, ct, vt)
- ι (standing user intent): the stable user intent, immutable
- ot (operational objective): the current actionable goal, revised based on evidence
- ct (constraints): known constraints, updated as discoveries are made
- vt (verification criteria): verification standards, adjusted as understanding deepens
- Candidate generation → accountability checks → authorized commit
- Failed branches can also become reusable progress
- Unverified narratives are never allowed to rewrite the entire task
- Token consumption is 1.41x, but mature waves save 21% solve tokens and 15% active time compared to startup waves
- Logged 34 verifier recoveries and 22 strict-review loop rescues
- An optimized RWKV6 kernel was merged upstream
- 640 campaign hours
- 254 bounded tasks
- 576 Engineer turns
- 286 Reviewer revisions
- 16 stage rollbacks
- All 6 pipelines reached submission completion
This separation lets the agent revise its operational path based on evidence without tampering with the original intent. The Manager holds change authority, but every change must be evidence-backed.
Verification-Gated Self-Evolution
Model weights never change. Evolution happens in runtime state — memory, skills, verifiers, routing policies — and only after role review and task-native verification can items be "committed."
Results
| Benchmark | Score | Comparison | |------|------|------| | SWE-Bench Pro | ~78% | Direct Copilot 59% | | AARRI-Bench | 76.8% | - | | Math data synthesis | +28.0 points | - |
Paper Pipeline Case Study
Six projects autonomously completed the paper production pipeline:
One-Sentence Summary
Argus doesn't make agents smarter — it makes agents not silently change goals, not fake completion, and hand intractable problems back to humans amid long-horizon uncertainty.