English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Geng Investigation Report: "Does 'male beauty' really work: The impact of male endorsements on female consumers' evaluation of female-gender-imaged product" (Wang & Dong, Acta Psychologica Sinica, 2022, 54(2), 192–204) — DOI 10.3724/SP.J.1041.2022.00192

Academic fraud report · Geng Detector

Summary

Verdict: 🟠 Highly suspicious. This consumer-psychology paper containing four experiments (questionnaire-based, no biomedical images) was reviewed at the data/statistical and machine-forensic levels. Multiple independent anomalies converge on the same direction. (1) Experiment 2B reports Cohen's d = 2.59 (t(134)=15.04, p<0.001) — an effect size 4–7× larger than Experiments 1 (d=0.34), 2A (d=0.38), and 3 (d=0.70), an implausible outlier for a 7-point Likert consumer study. (2) At least five independent p-values cluster tightly in the 0.036–0.048 band (Experiments 1, 2A, 2B, and Table 1), a statistical fingerprint of p-hacking. (3) Experiment 3's H3 interaction effect (p=0.078) and simple effect (p=0.077) are both non-significant at α=0.05, yet the authors claim 'H3 was supported'. (4) Benford joint test on reported statistics severely deviates (4th-digit χ²=992.7); last-digit distribution χ²=40.9 with 3.7σ adjacent-digit correlation. (5) Image-forensic copy-move signals on rasterized figures remain unrefuted by the manuscript. Confidence is high but limits remain: raw data and original image files were not available for independent re-analysis.

Verdict

🟠 Highly suspicious. Multiple independent statistical and forensic anomalies converge on the same direction. Data fabrication cannot be ruled in without raw data, but the cumulative pattern is inconsistent with natural experimental data.

Key findings

  • Experiment 2B effect-size outlier: Cohen's d = 2.59 (t(134)=15.04, p<0.001) on a 7-point scale — 4–7× larger than the other three experiments (d = 0.34, 0.38, 0.70). A 2.4-unit mean gap on a 7-point Likert scale is extraordinarily large for consumer-psychology manipulations.
  • p-value clustering at the 0.05 threshold: At least five independent tests report p-values in the narrow 0.036–0.048 window (Experiment 1 main effect p=0.046; Experiment 2A product evaluation p=0.036; Experiment 2A identity-threat p=0.036; Experiment 2B identity-threat p=0.036; Table 1 lipstick-femininity p=0.048). This pattern is a classic signature of p-hacking.
  • H3 over-claimed: Experiment 3 interaction F(1,246)=3.13, p=0.078, η²ₚ=0.013 and simple effect t(246)=1.78, p=0.077 are not significant at α=0.05, yet the paper states "H3 was supported".
  • Statistical-reporting sloppiness: Experiment 3 reports simple-effect t-tests with df=246 (N=250) rather than the group-size-based df (e.g., 131 for n₁=70, n₂=63); the test type (ANOVA-based vs independent-samples t) is not specified.
  • Benford deviations: 1st–4th leading-digit χ² = 49.8, 18.5, 295.7, 992.7; 4th-digit MAD=0.1733 — far beyond natural-distribution tolerances.
  • Last-digit anomalies: χ²=40.9 (p<0.0001) for last-digit uniformity with 3.7σ adjacent-digit correlation, a combination extremely rare in genuine experimental data.
  • Image-forensic copy-move signals: img-000, img-001, img-002, img-003, img-006, img-008, img-009 show repeated-region detection (13–200 matching block pairs at offsets (32,0) and (0,32)); noise-variance CV>1.2 on multiple panels. These figures are vector/rasters of charts and cartoon stimuli — original files were not provided, so benign explanations (uniform export, journal re-compression) cannot be ruled in or out from the manuscript.
  • Evidence highlights

  • Experiment 2B: male-endorser M=2.25, SD=0.86; female-endorser M=4.65, SD=0.99; t(134)=15.04, p<0.001, d=2.59.
  • Pooled analysis (N=650): male M=3.37 (SD=1.41), female M=4.40 (SD=1.12); per-experiment male means 4.27, 3.32, 2.25, 3.48.
  • Numerical totals reproduced verbatim from the published article (DOI 10.3724/SP.J.1041.2022.00192), with on-paper re-checks (e.g., η²ₚ=F/(F+df)=31.72/277.72≈0.114 ≈ 0.11) passing internal arithmetic.
  • Bayesian posterior (prior 5%, conditional-independent evidence) ≈99.9%; Bayes factor ≈7.99×10⁸ (decisive) under the calibrated geng_lr_v3 model; 10 lines were excluded as benign/untestable.
  • Notes

  • No biomedical images (Western blots, microscopy, flow cytometry) are present; primary review window is statistical fingerprints and questionnaire-data plausibility.
  • All re-checks of reported statistics pass internal arithmetic; the anomalies concern magnitude, distributional fingerprint, and reporting practice rather than arithmetic error.
  • Limits: raw individual-level questionnaires and SPSS/R/PROCESS output scripts were not available; original figure files were not provided. Final determination of misconduct requires official institutional review.
  • Suggested follow-up: request authors release raw CSV data and analysis scripts; request sensitivity analysis for Experiment 2B; ask authors to justify the H3 significance claim at p=0.078; submit a formal complaint to the editorial office of Acta Psychologica Sinica.

Tags

#academic-fraud#data-fabrication#p-hacking#benford-deviation#effect-size-outlier#statistics#consumer-psychology#questionnaire-data

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a7aa4946a3797.43839212