> Paper: DeGenTWeb: A First Look at LLM-dominant Websites > Authors: Sichang Steven He, Calvin Ardi, Ramesh Govindan, Harsha V. Madhyastha > arXiv: 2605.00087 | 2026-05-01
1. A New Era of "Content Farms"
You search for a solution to a technical problem. Google recommends a blog that looks professional—clear structure, perfect grammar, detailed information.
But you notice something odd:
- The "latest version" it mentions is actually two years old
- Some technical details sound plausible but don't hold up to scrutiny
- The author has no social media, no GitHub, no traceable identity
- Most of their content is LLM-generated
- Their purpose is usually SEO—harvesting traffic and ad revenue
- They often masquerade as human-written blogs, news, reviews, or Q&A
- These sites are growing rapidly
- They occupy top positions in some search query results
- Their content quality varies widely—some are nearly indistinguishable from human writing, others are riddled with factual errors
What you're reading may not be human-written. It may be LLM-generated.
2. LLMs Are "Colonizing" the Web
This paper presents the first systematic study of "LLM-dominant websites."
What are LLM-dominant websites?
Key findings:
3. Why Is Detection So Hard?
What's wrong with existing LLM text detectors?
1. High false-positive rates: To avoid mislabeling human content as AI-generated, detectors are tuned "loose"—letting large amounts of AI content slip through 2. Adversarial evolution: Content farms quickly learn to "bypass" detectors (fine-tuning, post-processing, mixing human and AI content) 3. Domain differences: Detectors trained on news may completely fail on technical blogs 4. Language bias: Detectors are less accurate on non-English content
This is an arms race—and the detectors are losing.
4. Impact: Degradation of the Information Ecosystem
What happens when LLM-generated content floods the web?
1. Training data pollution: Future LLMs will be trained on vast amounts of AI-generated text, leading to "model collapse" 2. Search devaluation: Finding genuinely valuable human content becomes harder and harder 3. Trust crisis: Readers can't judge whether content is reliable 4. Distorted economic incentives: Human creators are squeezed out of the market because AI content costs almost nothing
This isn't a problem of technological progress. It's a problem of the information ecosystem.
5. A Feynman-Style Judgment: Telling "Real" from "Real-Looking" Is an Old Human Problem
Feynman said:
> "Knowing the name of something" and "knowing something" are completely different.
The danger of LLM-generated content lies precisely here: it "looks like" knowledge, but it isn't. It is a probabilistic model recombining statistical patterns from training data.
When a website is full of content that "looks like" answers—but with no genuine understanding, no experience, no accountability—we are losing the ability to distinguish the "real" from the "real-looking."
6. Takeaways
As users and builders of the internet, ask yourself:
1. "How can I verify whether an information source is reliable?" 2. "How much of the online content I consume might be AI-generated?" 3. "If training data is polluted by AI-generated content, what will the next generation of AI become?" 4. "Do we need new 'digital signatures' to prove the human origin of content?"
DeGenTWeb is a warning. The value of the internet lies not in the quantity of content, but in its authenticity and diversity.
When AI starts generating training data for AI, we are heading toward an "echo chamber"—not of human echoes, but of machines talking to themselves.