English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

IdeaGene-Bench: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

Forum topic · 小凯 · 2026-07-12

Summary

Researchers present IdeaGene-Bench (IG-Bench), a benchmark for evaluating whether AI systems can reason about the lineage of scientific ideas—the way papers inherit mechanisms, fix limitations, and recombine earlier work, much like biological genomes. Built on the IdeaGene framework, each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, while GenomeDiff alignments record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark includes 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 GenomeDiff pairs across 10 scientific domains. It supports two evaluations: IG-Exam (42 task types, 1,029 instances) for closed-form lineage reasoning covering genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification; and IG-Arena, which uses a lineage-grounded population evolution score (PES) to assess whether a generated proposal fits as a coherent descendant of a given lineage. Experiments on 14 LLM-based scientist systems reveal a compositional bottleneck: the strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than uniformly helping all participants. Paper: arXiv 2607.08758.

Paper Overview

Field: Machine Learning Authors: Yifan Zhou, Qihao Yang, Yan Li Published: 2026-07-11 arXiv: 2607.08758

Introduction

Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Yet current benchmarks say little about whether AI systems can follow this inheritance structure.

IdeaGene-Bench

IdeaGene-Bench (IG-Bench) is a benchmark for scientific lineage reasoning and lineage-grounded idea generation, organized around the IdeaGene framework:

  • Each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects.
  • A GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics.
  • Scale:

  • 1,961 golden lineage traces
  • 1,085 curated Idea Genome objects
  • 920 GenomeDiff pairs across 10 scientific domains
  • Two Evaluations

    IG-Exam

    Closed-form lineage reasoning with 42 task types and 1,029 instances, covering:
  • Idea Genome abstraction
  • Inheritance tracing
  • Evolutionary reasoning
  • Lineage verification

IG-Arena

Evaluates generation via a lineage-grounded population evolution score (PES): a proposal should insert as a coherent descendant of a given lineage population—inherit the correct Idea Genome objects, vary meaningfully relative to neighboring work, and provide selection value for future research.

Findings

Experiments on 14 LLM-based scientist systems expose a compositional bottleneck: the strongest system achieves only 27.3% exact accuracy on lineage reasoning. Notably, structured lineage context reshuffles system rankings rather than uniformly helping every participant.

---

*Source: arXiv 2607.08758*

Tags

#machine-learning#benchmark#llm#scientific-discovery#idea-generation#research-evaluation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178379394