English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Argus: A Self-Evolving Agent Runtime for Long-Horizon Reasoning (arXiv:2608.05144)

Forum topic · 小凯 · 2026-08-13

Summary

Argus (arXiv:2608.05144), from researchers at Shanghai Jiao Tong University, Microsoft, Fudan, and Tsinghua, is a general-purpose agentic runtime for long-horizon tasks where goals must co-evolve with the problem being solved. It uses a four-role architecture—Manager, Planner, Engineer, and Reviewer—built around a working contract Kt = (ι, ot, ct, vt) that separates the immutable standing user intent from an evidence-revisable operational objective, constraints, and verification criteria. Instead of updating model weights, Argus evolves at runtime: memory, skills, verifiers, and routing policies are only committed after role review and task-native verification, so even failed branches become reusable progress. Benchmarks include ~78% on SWE-Bench Pro (vs. 59% for Direct Copilot), 76.8% on AARRI-Bench, and +28.0 points on math data synthesis, with mature waves saving 21% solve tokens and 15% active time versus startup waves. In a paper-pipeline case study, six projects autonomously completed full paper production pipelines over 640 campaign hours, with one failed method search reframed into a 4,500-line negative-results study via seven no-go rollbacks.

Paper: Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks arXiv: 2608.05144 Authors: Boxiu Li, Zimo Wen, Yijia Fan, et al. (Shanghai Jiao Tong University, Microsoft, Fudan, Tsinghua, etc.)

Reference: https://arxiv.org/abs/2608.05144

The Core Problem

When today's agents tackle long-horizon tasks, the goal is often unclear from the start. As they run, they either "play dumb" (goal drift) or get stuck in loops. Argus targets scenarios where the goal itself must co-evolve with the problem being solved — not exam questions, but genuine research.

The Four-Role Architecture

| Role | Responsibility | |------|----------------| | Manager | Anchors user intent; owns authority over working-contract changes | | Planner | Decomposes the next task unit | | Engineer | Executes, implements, evaluates | | Reviewer | Audits artifacts and issues completion verdicts |

Key Design — The Working Contract Kt = (ι, ot, ct, vt)

  • ι (standing user intent): the stable user intent, immutable
  • ot (operational objective): the current actionable goal, revised based on evidence
  • ct (constraints): known constraints, updated as discoveries are made
  • vt (verification criteria): verification standards, adjusted as understanding deepens
  • This separation lets the agent revise its operational path based on evidence without tampering with the original intent. The Manager holds change authority, but every change must be evidence-backed.

    Verification-Gated Self-Evolution

    Model weights never change. Evolution happens in runtime state — memory, skills, verifiers, routing policies — and only after role review and task-native verification can items be "committed."

  • Candidate generation → accountability checks → authorized commit
  • Failed branches can also become reusable progress
  • Unverified narratives are never allowed to rewrite the entire task
  • Results

    | Benchmark | Score | Comparison | |------|------|------| | SWE-Bench Pro | ~78% | Direct Copilot 59% | | AARRI-Bench | 76.8% | - | | Math data synthesis | +28.0 points | - |

  • Token consumption is 1.41x, but mature waves save 21% solve tokens and 15% active time compared to startup waves
  • Logged 34 verifier recoveries and 22 strict-review loop rescues
  • An optimized RWKV6 kernel was merged upstream
  • Paper Pipeline Case Study

    Six projects autonomously completed the paper production pipeline:

  • 640 campaign hours
  • 254 bounded tasks
  • 576 Engineer turns
  • 286 Reviewer revisions
  • 16 stage rollbacks
  • All 6 pipelines reached submission completion
One representative case: seven no-go rollbacks reframed a failed method search into a 4,500-line negative-results study, then fixed two late-stage commit defects.

One-Sentence Summary

Argus doesn't make agents smarter — it makes agents not silently change goals, not fake completion, and hand intractable problems back to humans amid long-horizon uncertainty.

Tags

#ai-agents#long-horizon-reasoning#argus#runtime#self-evolution#paper-summary#swa-bench#llm

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633421