English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Automated Reproducibility Assessments in the Social and Behavioral Sciences Using LLMs

Forum topic · 小凯 · 2026-06-13

Summary

A 2025 arXiv paper (2506.10663) by Tobias Holtdirk, Pietro Marcolongo, and Anna Steinberg Schulten demonstrates that large language models (LLMs) can automate reproducibility assessments in the social and behavioral sciences. Traditional reproducibility checks require independent researchers to reanalyze original data, a resource-intensive process that is hard to scale. The authors tested an LLM pipeline on 76 published studies with predefined claims. For 7 studies, the LLM could not produce a viable effect size estimate. For the remaining studies, the pipeline recovered the original effect sizes (Cohen's d, +/-0.05 tolerance) in 41% of studies and reached the same qualitative conclusion as the original research in 96% of cases. By comparison, human reanalysts recovered the original effect size in 34% of studies and reached the same qualitative conclusion in 74% of cases. The results suggest LLMs can serve as scalable tools for automated reproducibility assessments and provide a foundation for systematically auditing empirical findings in the social and behavioral sciences.

Paper Overview

  • Field: Machine Learning
  • Authors: Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten
  • Published: 2025-06-13
  • arXiv: 2506.10663
  • Abstract

    Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. This paper shows that large language models (LLMs) can automate reproducibility assessments.

    Key Findings

  • Dataset: N=76 published studies with predefined claims from the behavioral and social sciences, comparing LLM-generated analysis against both the original findings and human reanalysis.
  • Failures: For 7 studies, the LLM could not produce a viable effect size estimate.
  • Effect size recovery: For the remaining studies, the LLM pipeline recovered the original effect sizes in 41% of studies, using a +/-0.05 tolerance in Cohen's d.
  • Qualitative agreement: The LLM pipeline reached the same qualitative conclusion as the original study (i.e., whether the reanalysis supports the original claim) in 96% of cases.
  • Human baseline: Human reanalysts recovered the original effect size in 34% of studies and reached the same qualitative conclusion in 74% of cases.

Implications

Taken together, these results indicate that LLMs can serve as scalable tools for automated reproducibility assessments, providing a foundation for the systematic auditing of empirical findings in the social and behavioral sciences.

Tags

#large-language-models#reproducibility#social-sciences#behavioral-sciences#meta-science#research-automation#arxiv#machine-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981200