English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Integrity Review: Client evaluation decision models in the credit scoring tasks (Procedia Computer Science, 2020; DOI: 10.1016/j.procs.2020.09.068)

Academic fraud report · Geng Detector

Summary

Verdict: No evidence of academic fraud. This applied machine-learning conference paper by Ziemba, Radomska-Zalas, and Becker presents credit scoring models (logistic regression, optimized random forest, decision trees, etc.) evaluated with AUROC, Gini, and AUPRC. Six-style checks were applied where the paper type permitted. (1) Text appeared duplicated section-wide and Tables 3–4 appear twice, but this almost certainly reflects PDF-to-text artifacts from header/footer parsing or a two-up preview, not author misconduct; if present in the published PDF it is a layout defect only. (2) Reported Gini values are exactly consistent with the canonical relation Gini = 2 × AUROC − 1 across all spot-checked entries (e.g., 0.850 → 0.700; 0.881 → 0.762; 0.657 → 0.314), and AUPRC values are appropriately very low (0.013–0.276) for a highly imbalanced default dataset, indicating no fabricated metrics. (3) Feature selection was correctly performed on the training set only, ruling out common data-leakage inflation. (4) Hyperparameters (n_estimators=239, max_depth=13) suggest genuine grid-style tuning. Confidence in the clean verdict is moderate-to-high; the principal limitation is the impossibility of image forensics in a paper with no figures.

Verdict

No fraud detected. All numeric and methodological checks consistent with honest, applied ML research. Apparent text duplication is almost certainly a PDF parsing / typesetting artifact, not author misconduct.

Key findings

  • Numeric internal consistency (no fabricated metrics): Spot-checked AUROC↔Gini pairs satisfy Gini = 2·AUROC − 1 exactly; AUPRC values are appropriately low for a class-imbalanced default dataset.
  • No data-leakage pattern: Feature filtration was explicitly performed on the training set and only then applied to the test set (Section 3).
  • Plausible hyperparameter reporting: n_estimators = 239 and max_tree_depth = 13 are non-default, consistent with real grid-style tuning rather than copy-paste defaults.
  • 🟡 Layout/parsing artefact: Sections 1–5 and Tables 3–4 appear duplicated in the extracted text; this is most likely a PDF-to-text extraction error (header/footer bleed, two-up preview) rather than a substantive defect, but if present in the published file it would be a typesetting issue only.
  • N/A Image forensics: Paper contains no biological/chemical images (no Western blots, microscopy, etc.), so pixel-level reuse detection is not applicable.
  • Evidence highlights

  • DOI: 10.1016/j.procs.2020.09.068
  • Venue: Procedia Computer Science 176 (2020) 3301–3309; KES 2020.
  • Gini–AUROC cross-check (verbatim from tables):
  • Table 1, Logistic Regression: AUROC = 0.850 → Gini = 2(0.850) − 1 = 0.700 (matches table).
  • Table 3, Optimized Random Forest: AUROC = 0.881 → Gini = 0.762 (matches table).
  • Table 5, Decision Tree C4.5: AUROC = 0.657 → Gini = 0.314 (matches table).
  • Train/test hygiene: Author statement — "Feature filtration was performed on the training set, and the results were implemented into a set of test cases."
  • Class-imbalance realism: AUPRC (precision–recall for the positive class) ranges 0.013–0.276, consistent with a sparse-default credit dataset and inconsistent with the near-1.0 values typical of leakage-driven fraud.
  • Hyperparameters: number of iterations = 239, maximum tree depth = 13 (Optimized Random Forest).
  • Notes

  • Limits of this review: (i) no figures exist to inspect, so image-duplication checks were not performed; (ii) the apparent text duplication cannot be conclusively resolved without inspecting the publisher PDF directly; (iii) the paper uses a single 70/30 train/test split without k-fold cross-validation — a methodological thinness acceptable for Procedia-tier proceedings, not misconduct.
  • Recommended follow-up: none required with respect to author integrity; optionally verify the rendered PDF at Elsevier for the duplication artefact before citing.

Tags

#academic-integrity#credit-scoring#machine-learning#procedia-computer-science#methodology-check#numerical-consistency#data-leakage-check#no-figures

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/report/geng_geng_6a2e842f27a600.08225674