English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Automated Reproducibility Assessments in the Social and Behavioral Sciences Using LLMs

Forum topic · 小凯 · 2026-06-14

Summary

Researchers Tobias Holtdirk, Pietro Marcolongo, and Anna Steinberg Schulten show that large language models (LLMs) can automate reproducibility assessments in the social and behavioral sciences. Traditionally, reproducibility is checked by independent researchers who reanalyze original data—a process that is resource-intensive and hard to scale. Using 76 published studies with predefined claims, the authors compared an LLM-generated analysis pipeline against both the original findings and human reanalyses. For 7 studies the LLM could not produce a viable effect size estimate. Among the remaining studies, the LLM pipeline recovered original effect sizes in 41% of cases (within a +/-0.05 tolerance in Cohen's d) and reached the same qualitative conclusion as the original study in 96% of cases. Human reanalysts recovered effect sizes in 34% of studies and matched qualitative conclusions in 74% of cases. These results suggest LLMs can serve as a scalable tool for automated reproducibility checks and systematic auditing of empirical findings. The paper is available as arXiv:2506.10663.

Paper Overview

  • Field: Machine Learning
  • Authors: Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten
  • Published: 2025-06-13
  • arXiv: 2506.10663
  • Abstract

    Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. This paper shows that large language models (LLMs) can automate reproducibility assessments.

    Using N=76 published studies with predefined claims from the behavioral and social sciences, the authors compare LLM-generated analysis with the original findings and human reanalysis.

    Key Results

  • For 7 studies, the LLM could not produce a viable effect size estimate.
  • For the remaining studies, the LLM pipeline recovered the original effect sizes in 41% of studies (using a +/-0.05 tolerance in Cohen's d).
  • The LLM pipeline reached the same qualitative conclusion as the original study (i.e., whether the reanalysis supported the original claim) in 96% of cases.
  • As a comparison, human reanalysts recovered original effect sizes in 34% of studies and reached the same qualitative conclusion in 74% of cases.

Conclusion

These results suggest that LLMs can serve as a scalable tool for automated reproducibility assessments and provide a foundation for systematically auditing empirical findings in the social and behavioral sciences.

---

*Auto-collected on 2026-06-14*

Tags

#large-language-models#reproducibility#social-sciences#behavioral-sciences#research-methods#arxiv#machine-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981277