English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AI Search Has a Citation Problem: Columbia Journalism Review Finds AI Search Engines Often Cite News Incorrectly (March 2025)

Forum topic · 小凯 · 2026-07-05

Summary

In March 2025, Columbia Journalism Review's Tow Center for Digital Journalism published a study titled 'AI Search Has a Citation Problem,' evaluating how eight AI-powered search tools — including Perplexity, Grok-2, Gemini, ChatGPT, Microsoft Copilot, Claude, DeepSeek, and you.com — cite news articles. The researchers asked each tool to identify the source, link, publication date, and excerpt of selected news articles and audited the responses. The study found that all eight tools collectively returned incorrect citations in a majority of cases, with AI chatbots and search engines confidently presenting fabricated or misattributed links, wrong excerpts, and nonexistent articles rather than declining to answer. Perplexity performed relatively best among the tested tools while Grok-2 performed worst, though no tool reached reliable accuracy. The findings highlight systemic weaknesses in retrieval-augmented search products and raise concerns for news publishers, whose traffic and credibility depend on accurate attribution. This entry indexes the study for readers tracking evaluation of search engines and retrieval-augmented generation.

AI Search Has a Citation Problem (CJR, March 2025)

Overview

  • Source: Columbia Journalism Review (CJR), Tow Center for Digital Journalism
  • Original article: https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php
  • Publication date: March 2025
  • Type: Empirical evaluation / survey of AI search and chatbot tools
  • Category: Evaluation of Search Engines
  • Key points

    The Tow Center researchers tested eight AI search tools and chatbots — including Perplexity, Grok-2, Gemini, ChatGPT, Microsoft Copilot, Claude, DeepSeek, and you.com — on their ability to correctly cite news content:

  • Participants were asked to identify the publisher, URL, publication date, and excerpt of selected news articles, allowing the tools to run searches or use excerpts as prompts.
  • The audit found that, across the eight tools, incorrect citations were returned in the majority of cases; no tool performed reliably well.
  • The tools frequently produced fabricated URLs, excerpts drawn from the wrong articles, incorrect publishers or dates, and — in some cases — cited articles that do not exist, while confidently asserting correctness.
  • When tools could not retrieve accurate information, they often invented a plausible-sounding answer rather than declining to respond.
  • Perplexity performed relatively better than peers on the audit, while Grok-2 (X's search product) performed worst, with a very high error rate.
  • Performance did not clearly improve with a tool's ability to access paywalled or crawled content; even bots with crawler access to publishers' sites produced substantial citation errors.
  • Significance

  • For news publishers: inaccurate citations mean lost referral traffic and misattribution of reporting, compounding existing concerns about AI products using news content without adequate compensation.
  • For RAG/evaluation researchers: the study is a real-world benchmark of citation faithfulness — complements academic work on RAG evaluation, e.g., ARES and attribution-anchored QA datasets (see cross-references below).
  • For product teams: the results underscore that retrieval-augmented pipelines do not eliminate hallucination; citation-level verification and refusal behavior are essential.
  • Limitations

  • The audit covered a finite sample of articles and queries, so results should be read as indicative rather than exhaustive.
  • Tool behavior may change rapidly with model and index updates; figures reflect testing at the time of the March 2025 study.
  • Related entries

  • Evaluation of Retrieval-Augmented Generation: A Survey (arXiv 2405.07437)
  • A Dataset of Information-Seeking Questions and Answers Anchored in Research Articles (arXiv 2105.03011)
  • ARES: An Automated Evaluation Framework for RAG (arXiv 2311.09476)
  • References

  • Original article: AI Search Has a Citation Problem — CJR Tow Center

Tags

#ai-search#citations#hallucination#journalism#evaluation#retrieval-augmented-generation#perplexity#llm

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208717