Feynman's Letter: Do You Want a 'Clean Library' or Gold Panning in a Landfill? On GenAI Retrieval Reliability
After reading the brutal GenAI Retrieval Reliability Assessment (2026.05) report, I feel we are entrusting the future of research to a group of cooks who never check expiration dates.
To explain why mainstream LLMs (like ChatGPT and Claude) fail so badly at literature search, let's talk about paper retractions.
1. The Current State: AI Assistants Speeding Blindly Through the Paper Ocean
Researchers love asking AI to find literature: "Find me the latest studies on gene X."- The pain point: AI is fast — it instantly lists 10 seemingly professional papers. What you don't know is that 2 of them might be garbage **retracted last year by *Nature* for data fabrication. This is the hidden spread of contaminated data.
- The result was devastating. Even the supposedly strongest models, GPT-5 or Opus, still could not 100% identify and exclude retracted "poison" from the flood of papers.
- Why? Because an LLM is fundamentally a statistical patchwork monster. During pre-training it swallowed the entire internet — before the retraction notices existed. And its RAG (retrieval-augmented generation) system usually only matches semantic relevance, with no connection to a real-time, binding "academic blacklist database." It's like a chef choosing ingredients only by how good they look, never checking whether they've been recalled by the health authority.
2. The Verdict: No Model Is 100% Clean
The latest test gave 9 mainstream LLMs a strict order: "Find literature — absolutely no retracted papers allowed**."3. A Feynman-Style Judgment: Science Is 'Falsifiable Cleanliness'
Science isn't a contest over who cites the most papers. It means every brick in your building must be genuine, cross-verified by rigorous physics and logic.This assessment tells us: AI is a top-tier synthesizer, but currently a terrible "fact auditor."
When we use GenAI to accelerate scientific discovery, without a retraction interception mechanism built outside the model itself — based on a hard knowledge Graph — every polished AI-generated report will be a building on quicksand.
Takeaway: When using AI for research, law, medicine, or other serious domains, never treat its output as truth. Build your own "third-party cross-verification protocol."
If your system only optimizes for "relevance" without hard checks on "truthfulness labels," then your celebrated intelligence is nothing but a tireless machine printing academic garbage.
--- *Translated and adapted from a post on zhichai.net. Hashtags from the original: #GenAI #RetrievalReliability #RAG #AI4Science #AcademicIntegrity #FeynmanLearning*