Overview
Field: NLP Authors: Jieyuan Liu, Mengzhou Hu, Jefferson Chen, JungHo Kong, Pratibha Jagannatha, Yiming Gao, Dexter Pratt, Hsin-Yuan Lee, Zhiting Hu, Trey Ideker, Wei Wang, Eric P. Xing, Zhen Wang Published: 2026-09-14 arXiv: 2609.15938
Abstract
Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality remains an open question. Answering this question requires separating the effects of agents' scientific capabilities from those of their collaboration. A framework must therefore preserve agents' scientific roles and support rules for combining, revising, and retaining hypotheses.
Building on this view, the authors introduce HypoEvolve, which makes collaboration explicit through successive updates to a hypothesis population. Specifically, they propose a generational genetic algorithm to coordinate specialized LLM agents, integrating mechanistic argumentation, reconsideration of hypotheses, evidence assessment, and testability. Each generation specifies how scientific judgments and new proposals reshape the population, making the effect of collaboration on hypothesis quality directly testable.
Evaluation
The evaluation centers on scientifically meaningful hypotheses — explanations of how interventions might work. Drug repurposing links these explanations to target-level biological claims, assessed against external evidence. The authors adapt DepMap and Open Targets as complementary external metrics grounded in experimental, genetic, and clinical evidence.
Key results:
- Across 34 cancer types, HypoEvolve achieved the highest scores among six baselines on both metrics.
- DepMap selectivity reached 0.171, compared with 0.115 for the strongest baseline.
- Gains over single-shot generation generalized to held-out cancer types.
Significance
HypoEvolve advances a vision of autonomous science in which AI research teams achieve discovery capabilities beyond any single model.
--- *Auto-collected on 2026-09-16*