English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications

Forum topic · 小凯 · 2026-09-11

Summary

A research idea can be novel, coherent, and scientifically plausible while still being too underspecified for faithful implementation. The IdeaAMBIG paper studies the 'codification readiness' of implementation-facing research-method specifications: whether they provide enough methodological detail for a competent implementer or coding agent to build the intended method without unsupported assumptions. The authors build evidence-grounded specifications and supported resolutions from papers, codebases, issue threads, and reproduction artifacts, and introduce a benchmark of 660 evidence-grounded instances: 163 real-world gaps drawn from reproducibility reports and GitHub issues, plus 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Across 13 LLMs, the best model achieves only a 9.6% Macro Defect Recovery Rate on real-world instances, yet reaches an 80.6% Macro Clarification Action Success Rate when the defect is given. An oracle study shows that supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Defect localization is the main bottleneck for all evaluated models.

Paper Overview

  • Field: NLP
  • Authors: Yiling Ma, Yilun Zhao, Sihong Wu, Manasi Patwardhan, Arman Cohan
  • Published: 2026-09-09
  • arXiv: 2609.10539
  • Abstract

    A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. The paper studies the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. The authors construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts.

    The IdeaAMBIG Benchmark

    IdeaAMBIG contains 660 evidence-grounded instances:
  • 163 real-world gaps collected from reproducibility reports and GitHub issues
  • 497 controlled synthetic gaps injected into codification-ready references
  • The benchmark evaluates three capabilities: 1. Codification-readiness assessment — judging whether a specification is sufficiently complete 2. Defect localization — identifying underspecified points, given only the specification 3. Clarification action generation — producing resolutions, given the annotated defect

    Key Findings

  • Across 13 LLMs, the best model achieves only a 9.6% Macro Defect Recovery Rate on real-world instances.
  • Given the defect, the same best model reaches an 80.6% Macro Clarification Action Success Rate.
  • In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%.
  • Across all evaluated models, defect localization is the main bottleneck, with stronger performance on clarification once the defect is identified.
--- *Auto-collected on 2026-09-11*

Tags

#nlp#benchmark#llm#reproducibility#research-specifications#arxiv#coding-agents

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634705