Summary
This Chinese academic-integrity review (Geng) reports a high-suspicion verdict on the Nature paper "In vivo base editing of Chd3 rescues behavioural abnormalities in mice" (Yang, Li, Geng et al.; DOI: 10.1038/s41586-026-10113-6). The most serious finding concerns the non-human primate (NHP) experiments in Figure 5c, 5d, 5g: the legend states an untreated group of 2 monkeys and high- and low-dose groups of only 1 monkey each, yet the authors report unpaired two-sided t-tests yielding P < 0.0001 (****). With N=1 per group, standard deviation is undefined (df=0), making the t-test statistically meaningless; the paper appears to use tissue slices from a single animal as independent biological replicates (pseudoreplication). Additional concerns include p-values suspiciously clustered at the 0.05 threshold (P=0.0499 in Fig 1f; P=0.048 in Fig 4e; P=0.0435 in Fig 4g) suggestive of p-hacking, self-contradictory blinding statements, and the absence of formal normality testing before applying parametric tests to small samples (n=3–7). Limitations: this analysis is text-based only; no pixel-level image comparison was performed, so image-based duplication could not be assessed.
Verdict
🟠 Highly suspicious. The core NHP statistical analysis (Figure 5c, 5d, 5g) is methodologically untenable, and several smaller findings compound concerns about data handling. The author has stated that pixel-level image analysis was not possible in the current environment, so image duplication has not been assessed.
Key findings
- Pseudoreplication in NHP experiment (critical): According to the figure legend, the untreated group contains 2 monkeys, while the high-dose and low-dose groups each contain only 1 monkey. Despite N=1 per treatment group, the authors applied unpaired two-sided t-tests and reported P < 0.0001 (**) for Figure 5c, 5d, and 5g. With one animal per group, standard deviation is undefined and degrees of freedom = 0; any t-test result is statistically meaningless. The use of multiple brain slices (e.g., n=10 slices) from a single animal as independent biological replicates constitutes pseudoreplication.
- p-value clustering near 0.05 (highly suspicious): Multiple reported p-values sit suspiciously close to the 0.05 threshold, consistent with p-hacking:
- Figure 1f: P = 0.0499 (n = 6 independent experiments)
- Figure 4e: P = 0.048
- Figure 4g: P = 0.0435
- Methodological inconsistency on blinding: The Statistical analysis section states data collection/analysis were "not conducted in a blinded manner except for the immunohistochemical and behavioural tests…", yet the monkey immunohistochemistry methods section states "All joint laxity scores were recorded and analysed by investigators blinded to the genotypes of the mice." The two statements are internally contradictory, and non-blinded molecular data collection raises concerns about selective exclusion of "outlier" data points.
- No formal normality testing: The methods state "Data distribution was assumed to be normal, although this was not formally tested." Applying parametric tests (t-test, ANOVA) to small samples (n = 3, 4) without normality assessment is non-rigorous and can mask outliers.
- Template-style replication of prior work (concerning but not conclusive): The experimental design (TeABE, dual-AAV intracranial delivery, base editing in mouse brain) closely mirrors the authors' prior publication on Mef2c (ref 29: Nat. Neurosci. 27, 116–128 (2024)). The NHP ethics approval number Kmmu20205DS appears to date to 2020, used across a 4-year span for advanced in vivo base-editing work. This pattern suggests a "pipeline swap target" approach rather than independent data generation.
Evidence highlights
- DOI: 10.1038/s41586-026-10113-6
- NHP group sizes from Figure 5 legend: untreated n=2; high-dose n=1; low-dose n=1
- Statistical test applied with N=1: unpaired two-sided t-test
- Reported significance: P < 0.0001 (**) in Fig 5c, 5d, 5g
- Borderline p-values: 0.0499 (Fig 1f), 0.048 (Fig 4e), 0.0435 (Fig 4g)
- USV test sample size: n = 7 mice per group (Fig 2)
- Ethics approval: Kmmu20205DS (Kunming Medical University, 2020)
- Self-citation: ref 29, Nat. Neurosci. 27, 116–128 (2024), Mef2c base-editing study
Notes
- The reviewer explicitly notes that only text-based analysis was possible; high-resolution pixel comparison of figures was not conducted, so image manipulation or duplication cannot be ruled in or out from this report alone.
- The N=1 t-test issue is, on its face, sufficient to invalidate the NHP efficacy claims in Figure 5 and would, if confirmed by raw data, likely meet Nature's standards for a major correction or retraction.
- Other concerns (p-hacking, blinding inconsistency, lack of normality testing, template replication) are individually suggestive but not definitive; they strengthen the case for an institutional investigation and a request for individual-animal scatter plots rather than bar graphs.
- Recommended follow-up (per the original report): request raw NHP data and individual-level scatter plots; request the statistical code/software screenshots showing how the N=1 t-tests were executed; post the concern on PubPeer; submit a formal complaint to the Nature editorial board.
- This report is AI-assisted and intended for academic-discussion purposes only; a final determination of misconduct requires investigation by the authors' institution or the journal.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a33fd81b81e96.09697183