This post is a follow-up to an earlier piece on institutional problems with AI paper detection (Chinese original); this one focuses on the technology bubble itself—how absurd these tools really are, and what you can do about it.
1. The Absurd Experiment: Same Paper, 0% to 91%
One experiment is enough to destroy the credibility of AI detection tools.
Setup: Take one paper and run it through multiple mainstream AI detectors.
Result:
| Detection tool | "AI probability" for the same paper | |---|---| | Tool A | 0% (100% human) | | Tool B | 23% | | Tool C | 67% | | Tool D | 91% (almost certainly AI) |
Four tools, four wildly different verdicts on the same paper—a 91-percentage-point spread, from "definitely human" to "definitely AI."
> This is not error. This is a random number generator.
More ironic still: the paper is 100% human-written.
2. The "AI Premium" Paradox: The Better You Write, the More "AI" You Look
Detection tools have a counter-intuitive bias the academic literature calls the "AI Premium":
The higher your writing quality, the more likely you are flagged as AI-generated.
Why? Because the "high-quality text" in AI training data tends to be:
- Grammatically correct and clearly structured
- Logically tight with smooth transitions
- Free of typos and colloquialisms
- Formatted in standard academic style
- After AI writes it, have another AI (e.g., GPT-4) "rewrite this more casually and conversationally," or manually add a few typos and swap in unusual synonyms.
- Detection probability drops from ~90% to under 10%.
- AI writes the skeleton, humans fill in details; AI drafts, humans add personal anecdotes and concrete cases.
- Detectors are nearly helpless against mixed text—they cannot tell "which part is AI."
- Never generate the whole piece at once; generate in segments with human transitions and analysis in between.
- Detectors only output an "overall probability," making segmenting effective.
- OpenAI AI Text Classifier shutdown announcement (2023-07)
- "AI Detection Bias Against Non-Native English Writers" (2024, Stanford University)
- "The Impossibility of AI Detection" (2023, University of Maryland)
- Independent accuracy tests of Turnitin, GPTZero, iThenticate (multiple institutions, 2024–2025)
- "AI Premium" research: systematic misflagging of high-quality writing (2024, MIT/Stanford)
- Compilation of university AI-detection false-positive cases (student advocacy groups, 2025)
These are exactly the characteristics of strong student writing.
2.1 Anatomy of the False-Positive Mechanism
| Human trait misjudged | Detector's "logic" | Actual meaning | |---|---|---| | Correct grammar | "Too polished to be human" | Punishes well-trained writers | | Clear structure | "Too structured, like a template" | Punishes strong logical thinkers | | No typos | "Humans make mistakes, AI doesn't" | Punishes careful proofreaders | | Complex sentences | "High lexical diversity, like AI" | Punishes strong non-native speakers |
2.2 Who Gets Hurt
Research shows AI detectors are systematically biased against:
1. Non-native English-speaking scholars — they deliberately use more formal, "standard" phrasing to look academic. The harder they work at "correct English," the more AI-like they appear. Studies report false-positive rates 30–40% higher for non-native speakers. 2. Well-trained student writers — business, law, and medical writing programs teach "clear, concise, structured" prose, producing the highest false-positive rates. 3. STEM students — technical writing is naturally concise, formulaic, and low-variation, so CS, math, and engineering papers inherently resemble AI output patterns. 4. Anxious writers — stress makes writing more cautious, restrained, and "safe," which reads like "conservative AI."
3. Double Standards: AI Ghostwriters Bypass Detection; Honest Students Must Prove Innocence
3.1 How Easy Is It to Bypass Detection?
Method 1: "Humanize" the text
Method 2: Hybrid strategy
Method 3: Segmented generation
3.2 The Honest Student's Plight
| Scenario | What honest students face | |---|---| | Paper falsely flagged as AI | Hours or days spent writing a "self-defense report" | | The defense process | Providing drafts, notes, and thought processes—privacy exposure | | Financial cost | Some schools push "manual rewriting services" or "appeal services"—paying to prove innocence | | Psychological cost | Suspicion of misconduct brings stress, anxiety, even depression | | Time cost | Publication suspended, graduation delayed during appeals |
> Core contradiction: Those who use AI can easily bypass detection; those who don't must prove they didn't.
4. The Tech Bubble: Broken at the Foundation
4.1 Why AI Detection Is an Impossible Task
The basic assumption of AI detectors is: AI-generated text and human text have statistically distinguishable features.
But the assumption itself is flawed.
1. LLM training data = human text. GPT-4 was trained on trillions of tokens of human writing; its output distribution is a fit to the human writing distribution. The better the fit, the closer to human—and the harder to distinguish. 2. "High-quality human writing" ≈ "high-quality AI output." Their linguistic feature overlap exceeds 95%. Distinguishing them is like separating two drops of water with identical composition from different sources. 3. No stable "AI fingerprint." Different LLMs (GPT-4, Claude, Gemini, Llama) have different output characteristics; the same LLM varies with temperature settings. Detectors can only target known features of known models and fail on unknown ones.
4.2 OpenAI's Honesty
In July 2023, OpenAI shut down its own AI Text Classifier, officially admitting:
> "Our classifier has a low rate of accuracy and should not be used as a primary decision-making tool."
If the company that created LLMs admits it can't detect LLM output—why do third-party tools dare claim 98% accuracy?
4.3 The Business Model
| Stage | Move | Result | |---|---|---| | Manufacture panic | "AI ghostwriting is rampant—an academic integrity crisis!" | Universities and parents panic | | Sell the tool | "Our detector is 98% accurate" | Universities buy subscriptions | | Create demand | Detectors produce mass false positives → students need "rewriting services" | Detection companies offer "companion services" | | Harvest loop | Detect → false positive → appeal → paid service → re-detect | Continuous revenue |
This is not a conspiracy theory—it is reported commercial practice.
5. What You Can Do
5.1 If You're a Student
| Strategy | How | |---|---| | Transparent disclosure | State clearly which tools you used (Grammarly, ChatGPT polishing, etc.) | | Keep process evidence | Save drafts, revision history, notes—not to self-defend, but to protect your rights | | Refuse "detection as judgment" | If your school uses detectors, demand they publish accuracy and false-positive data | | Collective appeal | If falsely flagged, appeal together with other students—individual voices are weak, groups are strong | | Shift the burden of proof | Detector claims you used AI? It should prove it—not you proving innocence |
5.2 If You're an Educator / Reviewer
| Strategy | How | |---|---| | Don't use detectors | Refuse to use tools with < 90% accuracy as evaluation evidence | | Return to content evaluation | Ask "what problem does this paper solve," not "who wrote it" | | Oral defense | Have students explain core arguments, method choices, and failures | | Process transparency | Ask for "research logs"—not to catch cheating, but to understand the process | | Speak up | Voice opposition in your department to using detectors as hard criteria |
5.3 If You're an Administrator
| Strategy | How | |---|---| | Stop using detectors | Admit they don't work; stop wasting budget | | Reform assessment | Shift from "paper counts" to "problem solving" and "competency verification" | | Teach AI literacy | Teach students to use AI well, not to evade detection | | Protect students | Build appeal mechanisms so falsely flagged students have recourse |
5.4 If You're the Public
| Strategy | How | |---|---| | Spread the truth | Share detectors' real accuracy and false-positive rates | | Support victims | Follow the stories of falsely accused students and scholars | | Question the "integrity" narrative | When someone says "AI threatens academic integrity," ask what they mean by integrity | | Push institutional change | Advocate for assessment reform in academia and education |
6. Closing: Pop the Bubble, Return to Reality
AI detection tools are a proposition that fails at the technical foundation, packaged as "academic judges" to harvest university budgets and harm honest students.
But the deeper question is: why do we need a "judge" at all?
Because we can't evaluate "real research ability," we settle for evaluating "did you cheat." Because we can't measure "quality of thinking," we settle for measuring "does the text look like AI." Because we fear change, we use technology to prop up a collapsing evaluation system.
> Real academic integrity is not "didn't use AI"—it's "actually solved the problem."
Detection tools cannot detect real ability. They only detect "text features"—and text features have never been a proxy for ability.