English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AutoMat Benchmark: Can Coding Agents Reproduce Computational Materials Science Findings? Best Agents Reach Only 54.1% Success

Forum topic · 小凯 · 2026-05-05

Summary

Large language models perform strongly on software engineering benchmarks, but their ability to handle real computational science workflows remains unclear. This paper introduces AutoMat, a benchmark evaluating whether LLM-based coding agents can reproduce claims from real computational materials science papers. Built with domain experts, AutoMat poses three interrelated challenges: recovering underspecified computational procedures, navigating specialized toolchains, and judging whether resulting evidence supports a scientific claim. Evaluations across multiple coding agent setups and foundation models show low overall success, with the best setting achieving only 54.1%. Error analysis indicates agents perform worst when workflows must be reconstructed solely from paper text, failing mainly due to incomplete procedures, methodological deviations, and execution fragility. The authors, a team from institutions including Johns Hopkins researchers spanning NLP and materials science, position AutoMat as both a reproducibility benchmark and a diagnostic tool for agentic AI-for-science limitations. The paper (arXiv:2605.00803) highlights a significant gap between coding benchmark performance and genuine scientific workflow competence.

AutoMat: Evaluating Whether Coding Agents Can Reproduce Findings in Computational Materials Science

Paper Overview

  • Field: NLP / AI for Science
  • Authors: Ziyang Huang, Yi Cao, Ali K. Shargh, Jing Luo, Ruidong Mei, Mohd Zaki, Zhan Liu, Wyatt Bunstine, William Jurayj, Somdatta Goswami, Tyrel McQueen, Michael Shields, Jaafar El-Awady, Paulette Clancy, Benjamin Van Durme, Nicholas Andrews, William Walden, Daniel Khashabi
  • arXiv: 2605.00803
  • Background

    Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims.

    AutoMat Benchmark

    To address this question, the authors present AutoMat, a benchmark for evaluating LLM-based agents' ability to reproduce claims from computational materials science. AutoMat poses three interrelated challenges:

    1. Recovering underspecified computational procedures — many papers leave key computational steps implicit. 2. Navigating specialized toolchains — domain-specific software and workflows that agents must orchestrate. 3. Evidence evaluation — determining whether the resulting evidence supports a scientific claim.

    By working closely with subject matter experts, the team curated a set of claims from real materials science papers to test whether coding agents can recover and execute the end-to-end workflow needed to support (or undermine) such claims.

    Key Findings

  • Current LLM-based agents obtain low overall success rates on AutoMat, with the best-performing setting achieving a success rate of only 54.1%.
  • Agents perform worst when workflows must be reconstructed from paper text alone.
  • Failures stem primarily from:
  • Incomplete procedures
  • Methodological deviations
  • Execution fragility

Significance

Taken together, these findings position AutoMat as both a benchmark for computational scientific reproducibility and a tool for diagnosing the current limitations of agentic systems in AI-for-science settings.

Abstract (Original)

> Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims. To address this question, we present AutoMat, a benchmark for evaluating LLM-based agents' ability to reproduce claims from computational materials science. AutoMat poses three interrelated challenges: recovering underspecified computational procedures, navigating specialized toolchains, and determining whether the resulting evidence supports a claim. By working closely with subject matter experts, we curate a set of claims from real materials science papers to test whether coding agents can recover and execute the end-to-end workflow needed to support (or undermine) such claims. We then evaluate multiple representative coding agent settings across several foundation models. Our results show that current LLM-based agents obtain low overall success rates on AutoMat, with the best-performing setting achieving a success rate of only 54.1%. Error analysis further reveals that agents perform worst when workflows must be reconstructed from paper text alone and that they fail primarily due to incomplete procedures, methodological deviations, and execution fragility. Taken together, these findings position AutoMat as both a benchmark for computational scientific reproducibility and a tool for diagnosing the current limitations of agentic systems in AI-for-science settings.

*Auto-collected on 2026-05-05*

Tags

#llm-agents#benchmark#ai-for-science#materials-science#reproducibility#nlp#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619469