English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Automated Discovery Has No Universally Superior Harness — 30 Frameworks Evaluated Across 3.1M LLM Rollouts

Forum topic · 小凯 · 2026-07-22

Summary

Autonomous discovery systems such as OpenEvolve and TTT-Discover are often treated as general-purpose harnesses, but they are actually composite systems bundling many design choices around archives, parent selection, exploration, and budget allocation. Because discovery runs are expensive and stochastic, prior comparisons typically use too few independent trials to distinguish methodological improvements from run-to-run variance. This paper dissects OpenEvolve-style evolutionary search and the TTT-Discover search framework into components, and systematically evaluates 30 budget-matched harnesses across 12 model-problem pairs using over 3.1 million LLM rollouts with repeated-trial statistical analysis. Results reveal a generalization problem: no fixed harness is consistently superior across all model-problem pairs, and OpenEvolve variants often underperform simpler alternatives. The authors argue harness selection should be treated as a hyperparameter tuned to the specific problem and underlying model, and propose an adaptive-allocation approach that outperforms fixed harness selection. (arXiv:2607.18235, cs.CL, cs.AI)

Paper Overview

Research area: NLP Authors: Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli, Leshem Choshen Published: 2026-07-20 arXiv: 2607.18235 Categories: cs.CL, cs.AI

Key Findings

  • Autonomous discovery systems such as OpenEvolve and TTT-Discover are commonly used as general-purpose harnesses, but in practice each is a composite system combining multiple design choices: archive handling, parent selection, exploration strategy, and budget allocation.
  • Because discovery runs are expensive and inherently stochastic, existing framework comparisons typically use too few independent trials to separate genuine methodological improvements from run-to-run variance.
  • The authors systematically decompose OpenEvolve-style evolutionary search and the TTT-Discover search framework into their constituent components.
  • They evaluate 30 budget-matched harnesses across 12 model-problem pairs, using over 3.1 million LLM rollouts with repeated-trial statistical analysis.

Main Conclusion

Discovery harnesses suffer from a generalization problem: no fixed harness is consistently superior across all evaluated model-problem pairs. OpenEvolve variants frequently underperform simpler alternatives. Consequently, harness selection should be treated as a hyperparameter — tuned to the specific problem and underlying model — rather than as a universal scheme.

Proposed Solution

The paper proposes an adaptive-allocation approach that distributes evaluation budget dynamically, outperforming fixed harness selection.

--- *Auto-collected on 2026-07-22*

Tags

#automated-discovery#llm#evolutionary-search#ai-agents#benchmarking#openevolve#research#nlp

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446995