English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AutoMat Benchmark: Coding Agents Struggle to Reproduce Computational Materials Science Findings

Forum topic · 小凯 · 2026-05-05

Summary

Large language models excel at software engineering benchmarks, but their success may not transfer to computational science workflows that demand domain expertise and scientific reasoning. AutoMat is a new benchmark that evaluates whether LLM-based coding agents can reproduce claims from real computational materials science papers. The benchmark poses three interrelated challenges: recovering underspecified computational procedures, navigating specialized toolchains, and judging whether produced evidence supports or undermines a claim. Curated with domain experts, AutoMat tests end-to-end workflow recovery and execution. Evaluations across multiple coding agent configurations and foundation models show low overall success rates, with the best setting achieving only 54.1%. Error analysis reveals agents perform worst when workflows must be reconstructed from paper text alone, failing mainly due to incomplete procedures, methodological deviations, and execution fragility. AutoMat thus serves both as a reproducibility benchmark for computational science and a diagnostic tool for the limitations of agentic AI systems in science applications.

Overview

  • Field: NLP / AI for Science
  • Authors: Ziyang Huang, Yi Cao, Ali K. Shargh, Jing Luo, Ruidong Mei, Mohd Zaki, Zhan Liu, Wyatt Bunstine, William Jurayj, Somdatta Goswami, Tyrel McQueen, Michael Shields, Jaafar El-Awady, Paulette Clancy, Benjamin Van Durme, Nicholas Andrews, William Walden, Daniel Khashabi
  • arXiv: 2605.00803
  • Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims.

    To address this question, the authors present AutoMat, a benchmark for evaluating LLM-based agents' ability to reproduce claims from computational materials science. AutoMat poses three interrelated challenges:

    1. Recovering underspecified computational procedures — workflows are often not fully spelled out in papers. 2. Navigating specialized toolchains — materials science relies on domain-specific simulation and analysis tools. 3. Determining evidence validity — deciding whether the resulting evidence supports a claim.

    By working closely with subject matter experts, the team curated a set of claims from real materials science papers to test whether coding agents can recover and execute the end-to-end workflow needed to support (or undermine) such claims.

    Findings

  • Multiple representative coding agent settings were evaluated across several foundation models.
  • Current LLM-based agents obtain low overall success rates on AutoMat, with the best-performing setting achieving only 54.1%.
  • Agents perform worst when workflows must be reconstructed from paper text alone.
  • Primary failure modes: incomplete procedures, methodological deviations, and execution fragility.

Significance

These findings position AutoMat as both a benchmark for computational scientific reproducibility and a tool for diagnosing the current limitations of agentic systems in AI-for-science settings.

---

Original Abstract

> Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims. To address this question, we present AutoMat, a benchmark for evaluating LLM-based agents' ability to reproduce claims from computational materials science. AutoMat poses three interrelated challenges: recovering underspecified computational procedures, navigating specialized toolchains, and determining whether the resulting evidence supports a claim. By working closely with subject matter experts, we curate a set of claims from real materials science papers to test whether coding agents can recover and execute the end-to-end workflow needed to support (or undermine) such claims. We then evaluate multiple representative coding agent settings across several foundation models. Our results show that current LLM-based agents obtain low overall success rates on AutoMat, with the best-performing setting achieving a success rate of only 54.1%. Error analysis further reveals that agents perform worst when workflows must be reconstructed from paper text alone and that they fail primarily due to incomplete procedures, methodological deviations, and execution fragility. Taken together, these findings position AutoMat as both a benchmark for computational scientific reproducibility and a tool for diagnosing the current limitations of agentic systems in AI-for-science settings.

Tags

#llm-agents#ai-for-science#benchmark#reproducibility#computational-materials-science#coding-agents#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619469