Paper Overview
- Field: Machine Learning
- Authors: Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten
- Published: 2025-06-13
- arXiv: 2506.10663
- For 7 studies, the LLM could not produce a viable effect size estimate.
- For the remaining studies, the LLM pipeline recovered the original effect sizes in 41% of studies (using a +/-0.05 tolerance in Cohen's d).
- The LLM pipeline reached the same qualitative conclusion as the original study (i.e., whether the reanalysis supported the original claim) in 96% of cases.
- As a comparison, human reanalysts recovered original effect sizes in 34% of studies and reached the same qualitative conclusion in 74% of cases.
Abstract
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. This paper shows that large language models (LLMs) can automate reproducibility assessments.
Using N=76 published studies with predefined claims from the behavioral and social sciences, the authors compare LLM-generated analysis with the original findings and human reanalysis.
Key Results
Conclusion
These results suggest that LLMs can serve as a scalable tool for automated reproducibility assessments and provide a foundation for systematically auditing empirical findings in the social and behavioral sciences.
---
*Auto-collected on 2026-06-14*