Paper Overview
- Field: Machine Learning
- Authors: Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten
- Published: 2025-06-13
- arXiv: 2506.10663
- Dataset: N=76 published studies with predefined claims from the behavioral and social sciences, comparing LLM-generated analysis against both the original findings and human reanalysis.
- Failures: For 7 studies, the LLM could not produce a viable effect size estimate.
- Effect size recovery: For the remaining studies, the LLM pipeline recovered the original effect sizes in 41% of studies, using a +/-0.05 tolerance in Cohen's d.
- Qualitative agreement: The LLM pipeline reached the same qualitative conclusion as the original study (i.e., whether the reanalysis supports the original claim) in 96% of cases.
- Human baseline: Human reanalysts recovered the original effect size in 34% of studies and reached the same qualitative conclusion in 74% of cases.
Abstract
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. This paper shows that large language models (LLMs) can automate reproducibility assessments.
Key Findings
Implications
Taken together, these results indicate that LLMs can serve as scalable tools for automated reproducibility assessments, providing a foundation for the systematic auditing of empirical findings in the social and behavioral sciences.