Summary
AI4AI-Bench is a new benchmark for evaluating whether LLM agents can improve the training algorithms that produce AI systems—the core of recursive self-improvement (RSI). Existing benchmarks cannot isolate this capability, so the authors built a suite of 10 frozen research repositories spanning 10 training algorithm families. In each task, an agent has 4 hours on a single B300 GPU to rewrite a training algorithm; the resulting code is then rerun from scratch for up to 12 hours and scored by a fixed evaluator. Because the 10 task metrics are incomparable, each is normalized: 0 represents an uninformative model, 0.1 the repository's original algorithm, and 1.0 the task optimum. Across 29 configurations from 6 systems, the average score is 0.166, with the best system reaching 0.250—meaning even the strongest closes less than a fifth of the gap between existing and optimal algorithms. Most submissions never changed how the model learns; the few that did averaged 0.226 versus 0.126 for the rest. More inference effort primarily increases willingness to attempt algorithmic change. arXiv: 2608.20318.
Paper Overview
Field: NLP
Authors: Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang
Published: 2026-08-22
arXiv: 2608.20318
Abstract
Recursive self-improvement (RSI) asks whether AI systems can improve the process that produces AI systems. That process is the training algorithm: better objectives or update rules improve the compute-capability exchange rate of each subsequent run—including runs that produce the next agent. Existing benchmarks cannot isolate this capability.
This paper proposes AI4AI-Bench, consisting of 10 frozen research repositories covering 10 training algorithm families. In each task, the agent has 4 hours on a single B300 to rewrite a training algorithm; its code is then rerun from scratch for up to 12 hours and scored by a fixed evaluator.
Because the 10 metrics are incomparable, each task is mapped to a common scale: 0 for an uninformative model, 0.1 for the repository's original algorithm, and 1.0 for the task optimum.
Key Results
- 29 configurations from 6 systems average 0.166 across all 10 tasks; the best system reaches 0.250—even the strongest closes less than one-fifth of the gap between the existing algorithm and the optimum.
- Most submissions never changed how the model learns; the few that did averaged 0.226, versus 0.126 for the rest.
- More inference effort mainly buys the willingness to attempt algorithmic change, not better changes.
---
*Auto-collected on 2026-08-22*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178633815