Paper Overview
- Field: Machine Learning (ML)
- Authors: Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten, Felix Henninger, Stefan Rose, Sarah Ball, Bolei Ma, Frauke Kreuter, Markus Weinmann, Stefan Feuerriegel
- Published: 2026-06-11
- arXiv: 2606.13670
- For 7 studies, the LLM could not produce a viable effect size estimate.
- For the remaining studies, the LLM pipeline recovered the original effect sizes in 41% of studies using a +/-0.05 tolerance in Cohen's d.
- The LLM pipeline reached the same qualitative conclusion as the original study in 96% of cases (i.e., whether the reanalysis supported the original claim).
- In comparison, human reanalysts recovered the original effect sizes in 34% of studies and reached the same qualitative conclusion in 74% of cases.
Abstract
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. Here, the authors show that large language models (LLMs) can automate reproducibility assessments.
Using N=76 published studies with predefined claims from the behavioral and social sciences, they compare LLM-generated analysis with the original findings and human reanalysis.
Key Findings
Implications
These results suggest that LLMs can serve as a scalable tool for automated reproducibility assessments, providing a foundation for systematically auditing empirical findings in the social and behavioral sciences.
--- *Auto-collected on 2026-06-15*