English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications

Forum topic · 小凯 · 2026-09-11

Summary

A research idea can be novel, coherent, and scientifically plausible while still being underspecified for faithful implementation. The IdeaAMBIG benchmark (arXiv:2609.10539) by Yiling Ma, Yilun Zhao, Sihong Wu, Manasi Patwardhan, and Arman Cohan studies the codification readiness of implementation-facing research-method specifications: whether they contain enough methodological detail for an implementer or coding agent to build the intended method without unsupported assumptions. The benchmark comprises 660 evidence-grounded instances drawn from papers, codebases, issue threads, and reproduction artifacts: 163 real-world gaps sourced from reproducibility reports and GitHub issues, plus 497 controlled synthetic gaps injected into codification-ready references. It evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Across 13 LLMs, the best model achieves only a 9.6% Macro Defect Recovery Rate on real-world instances, but an 80.6% Macro Clarification Action Success Rate when the defect is given, identifying defect localization as the primary bottleneck. An oracle study shows that providing the gold resolution raises the downstream codification-ready rate from 14% to 98%.

Overview

Research area: NLP Authors: Yiling Ma, Yilun Zhao, Sihong Wu, Manasi Patwardhan, Arman Cohan Published: 2026-09-09 arXiv: 2609.10539

Abstract

A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. This work studies the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions.

The authors construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts.

The IdeaAMBIG Benchmark

IdeaAMBIG contains 660 evidence-grounded instances:

  • 163 real-world gaps collected from reproducibility reports and GitHub issues
  • 497 controlled synthetic gaps injected into codification-ready references
  • The benchmark evaluates three capabilities:

    1. Codification-readiness assessment 2. Defect localization (given only the specification) 3. Clarification action generation (given the specification plus the annotated defect)

    Key Findings

  • Across 13 LLMs, the best model achieves only a 9.6% Macro Defect Recovery Rate on real-world instances, but an 80.6% Macro Clarification Action Success Rate when the defect is provided.
  • Defect localization is the main bottleneck across all evaluated models; models are much stronger at generating clarifications once a defect is identified.
  • In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%.
  • Resources

  • Paper: <https://arxiv.org/abs/2609.10539>
---

*Auto-collected on 2026-09-11.*

Tags

#nlp#llm-benchmark#reproducibility#research-ideas#coding-agents#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634716