English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Forum topic · 小凯 · 2026-09-23

Summary

RRSI (Regularized Recursive Self-Improvement of Agent Harnesses) addresses overfitting in automated agent harness evolution. LLM agent capability depends heavily on the harness — prompts, control flow, tooling, memory, and context management around a frozen backbone model. Recent methods automate harness improvement via iterative component-wise edits, a form of recursive self-improvement (RSI), but these often memorize training tasks: large in-distribution gains shrink or vanish out-of-distribution. RRSI introduces regularization into harness self-improvement by constraining both proposal and selection of evolution candidates. The proposer uses a time-annealed budget limiting how many edits can be bundled, encouraging exploration of untaken paths based on evolution history. The selector employs a critic to filter benchmark-specific proposals and a pruner to remove changes that are too small, too expensive, or no longer useful. Across eight benchmarks spanning coding, agentic workspace, and engineering design tasks, RRSI improves up to 14.1 points on targeted splits and up to 4.7 points on five out-of-distribution benchmarks, while producing harnesses that consume 30% less policy tokens than unregularized evolution. Code is open-sourced.

Paper Overview

Field: NLP Authors: Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhang, Yoonho Lee, Chengsong Huang, Han Yu, Zhongying CuiZhu, Yifei Ming, Huaxiu Yao, Burak Gokturk, Tomas Pfister, Chen-Yu Lee Published: 2026-09-21 arXiv: 2609.24972

Background

An LLM agent's capability is largely magnified by its *harness* — the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level.

The Problem

Such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks.

The RRSI Method

RRSI incorporates the principles of regularization into harness self-improvement by constraining the evolution candidate proposal and selection process:

  • Proposer: uses a time-annealed budget that limits the number of edits candidates can bundle, and encourages exploration of untaken paths based on evolution history.
  • Selector: is equipped with a *critic* that filters benchmark-specific proposals, and a *pruner* that removes changes that are too small, too expensive, or no longer useful.
  • Together, these constraints favor reusable agent mechanisms over benchmark-specific ones or noise.

    Results

    Across eight benchmarks spanning coding, agentic workspace, and engineering design tasks:

  • Up to 14.1 points improvement on splits targeted by evolution
  • Up to 4.7 points improvement on five out-of-distribution benchmarks
  • Harnesses consume 30% less policy tokens than unregularized evolution

Resources

Code: https://github.com/google-research/rrsi

---

*Auto-collected on 2026-09-23*

Tags

#rrsi#recursive-self-improvement#llm-agents#regularization#harness-evolution#nlp#arxiv#out-of-distribution-generalization

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635104