Summary
This report examines the Nature article by Romera-Paredes and colleagues from Google DeepMind describing FunSearch, a system combining large language models with evolutionary search to discover new results in combinatorial mathematics and online bin packing. The verdict is clean (✅). As a computer science / discrete mathematics paper, the work contains no wet-lab biology figures and shows no signs of image duplication, splicing, or pixel-level manipulation. The authors openly report a low success rate on the hardest Cap Set problem (n=8): only 4 of 140 independent runs produced a size-512 cap set, which is the kind of honest negative result typically absent in fabricated data. Numerical results in Table 1 (e.g., FunSearch excess-bin fractions 5.30%, 4.19%, 3.11%, 2.47% across OR1–OR4) and the discrete-math improvement from 496 to 512 display natural, non-linear variation rather than fabricated regularity. The timeline is consistent: submitted 12 Aug 2023, accepted 30 Nov 2023, uses the PaLM-2-based Codey model released in 2023. Code is released at github.com/google-deepmind/funsearch. Limitations: statistical and visual checks are mostly inapplicable to this domain, and final judgement of fabrication requires institutional investigation.
Verdict
✅
Clear (no fraud indicators detected). The paper shows no hallmarks of image manipulation, data fabrication, statistical irregularities, timeline inconsistencies, or methodological contradictions. The verdict is limited by the fact that most standard fraud-detection modalities (image forensics, p-value irregularities) do not meaningfully apply to a code-and-mathematics paper; conclusions therefore rest on numerical sanity checks, transparency around negative results, timeline plausibility, and open code release.
Key findings
- No figures rely on wet-lab imaging; the paper uses algorithm diagrams, code excerpts, and data tables. No pixel duplication, mirroring, splicing, or edge-clipping signatures were found in Figures 1–6 or Table 1.
- The Cap Set (n=8) experiment is reported with unusually high transparency: 140 independent runs produced only 4 size-512 cap sets, contradicting the pattern in which fabricated results overstate success.
- Reported numerical gains (e.g., improvement from 496 to 512 in cap set size; FunSearch excess-bin fractions 5.30%, 4.19%, 3.11%, 2.47% on OR1–OR4) show natural, non-linear variation inconsistent with synthetic random patterns (no uniform trailing digits, no identical standard deviations).
- Timeline is internally consistent: submitted 2023-08-12, accepted 2023-11-30; uses the Codey/PaLM-2 model family released by Google mid-2023 with available API access. References are bounded by 2023 preprints, including GPT-4.
- Methods are described in unusually fine engineering detail (distributed system, island-based evolutionary model, e.g., discarding the worst-performing half of islands every 4 hours), with no internal contradictions.
- Code is released at
https://github.com/google-deepmind/funsearch; the repository is live and consistent with the paper's description. Evidence highlights
- Cap Set success-rate disclosure (Methods): "only 4 of 140 independent runs" produced a size-512 cap set — a signature of authentic stochastic search rather than cherry-picking.
- Table 1, online bin packing: FunSearch excess-bin fractions 5.30%, 4.19%, 3.11%, 2.47% across instances OR1–OR4; curve is monotonically decreasing and non-arithmetic, consistent with real optimization output.
- Discrete-math claim: cap-set construction improved from 496 to 512 — a verifiable combinatorial result, not an arbitrary metric.
- DOI: 10.1038/s41586-023-06924-6, Nature Vol. 625, online 14 December 2023.
- Code availability: https://github.com/google-deepmind/funsearch
Notes
- Standard image and p-value forensics are largely N/A for this contribution; the clean verdict should be read as "no positive indicators of fraud," not "proof of innocence."
- Independent reproduction of FunSearch on comparable compute would be the strongest confirmation; the authors have facilitated this by open-sourcing the code.
- Final adjudication of any academic-integrity concern remains the responsibility of the publishing journal and/or institutional review bodies.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/report/geng_geng_6a2eed5a5df8e9.85668589