AI Search Has a Citation Problem (CJR, March 2025)
Overview
- Source: Columbia Journalism Review (CJR), Tow Center for Digital Journalism
- Original article: https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php
- Publication date: March 2025
- Type: Empirical evaluation / survey of AI search and chatbot tools
- Category: Evaluation of Search Engines
- Participants were asked to identify the publisher, URL, publication date, and excerpt of selected news articles, allowing the tools to run searches or use excerpts as prompts.
- The audit found that, across the eight tools, incorrect citations were returned in the majority of cases; no tool performed reliably well.
- The tools frequently produced fabricated URLs, excerpts drawn from the wrong articles, incorrect publishers or dates, and — in some cases — cited articles that do not exist, while confidently asserting correctness.
- When tools could not retrieve accurate information, they often invented a plausible-sounding answer rather than declining to respond.
- Perplexity performed relatively better than peers on the audit, while Grok-2 (X's search product) performed worst, with a very high error rate.
- Performance did not clearly improve with a tool's ability to access paywalled or crawled content; even bots with crawler access to publishers' sites produced substantial citation errors.
- For news publishers: inaccurate citations mean lost referral traffic and misattribution of reporting, compounding existing concerns about AI products using news content without adequate compensation.
- For RAG/evaluation researchers: the study is a real-world benchmark of citation faithfulness — complements academic work on RAG evaluation, e.g., ARES and attribution-anchored QA datasets (see cross-references below).
- For product teams: the results underscore that retrieval-augmented pipelines do not eliminate hallucination; citation-level verification and refusal behavior are essential.
- The audit covered a finite sample of articles and queries, so results should be read as indicative rather than exhaustive.
- Tool behavior may change rapidly with model and index updates; figures reflect testing at the time of the March 2025 study.
- Evaluation of Retrieval-Augmented Generation: A Survey (arXiv 2405.07437)
- A Dataset of Information-Seeking Questions and Answers Anchored in Research Articles (arXiv 2105.03011)
- ARES: An Automated Evaluation Framework for RAG (arXiv 2311.09476)
- Original article: AI Search Has a Citation Problem — CJR Tow Center
Key points
The Tow Center researchers tested eight AI search tools and chatbots — including Perplexity, Grok-2, Gemini, ChatGPT, Microsoft Copilot, Claude, DeepSeek, and you.com — on their ability to correctly cite news content: