English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Automating Reproducibility Assessments in the Social and Behavioral Sciences with LLMs

Forum topic · 小凯 · 2026-06-15

Summary

A new arXiv paper (2606.13670) by Tobias Holtdirk, Pietro Marcolongo, and colleagues including Stefan Feuerriegel shows that large language models (LLMs) can automate reproducibility assessments in the social and behavioral sciences. Traditional reproducibility checks rely on independent human reanalyses of original data, which are resource-intensive and hard to scale. The authors evaluated an LLM pipeline on 76 published studies with predefined claims, comparing LLM-generated analyses against original findings and human reanalysis. For 7 studies, the LLM could not produce a viable effect size estimate. Among the remaining studies, the pipeline recovered original effect sizes in 41% of cases using a ±0.05 tolerance in Cohen's d, and reached the same qualitative conclusion as the original study in 96% of cases. Human reanalysts, by comparison, recovered original effect sizes in 34% of studies and matched qualitative conclusions in 74% of cases. These results suggest LLMs can serve as a scalable tool for automated reproducibility assessment and enable systematic auditing of empirical findings across the social and behavioral sciences.

Paper Overview

  • Field: Machine Learning (ML)
  • Authors: Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten, Felix Henninger, Stefan Rose, Sarah Ball, Bolei Ma, Frauke Kreuter, Markus Weinmann, Stefan Feuerriegel
  • Published: 2026-06-11
  • arXiv: 2606.13670
  • Abstract

    Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. Here, the authors show that large language models (LLMs) can automate reproducibility assessments.

    Using N=76 published studies with predefined claims from the behavioral and social sciences, they compare LLM-generated analysis with the original findings and human reanalysis.

    Key Findings

  • For 7 studies, the LLM could not produce a viable effect size estimate.
  • For the remaining studies, the LLM pipeline recovered the original effect sizes in 41% of studies using a +/-0.05 tolerance in Cohen's d.
  • The LLM pipeline reached the same qualitative conclusion as the original study in 96% of cases (i.e., whether the reanalysis supported the original claim).
  • In comparison, human reanalysts recovered the original effect sizes in 34% of studies and reached the same qualitative conclusion in 74% of cases.

Implications

These results suggest that LLMs can serve as a scalable tool for automated reproducibility assessments, providing a foundation for systematically auditing empirical findings in the social and behavioral sciences.

--- *Auto-collected on 2026-06-15*

Tags

#llms#reproducibility#machine-learning#social-sciences#research-methods#arxiv#automated-analysis

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981343