C-ReD: A Comprehensive Chinese Benchmark for AI-Generated Text Detection Derived from Real-World Prompts
Paper Overview
- Research areas: cs.CL, cs.AI
- Authors: Chenxi Qing, Junxi Wu, Zheng Liu, Yixiang Qiu, Hongyao Yu, Bin Chen, Hao Wu, Shu-Tao Xia
- Published: 2026-04-13
- arXiv: 2604.11796
Summary (translated)
Recent large language models (LLMs) are capable of generating highly fluent textual content. While they offer significant convenience to humans, they also introduce various risks, such as phishing and academic dishonesty. Numerous research efforts have been dedicated to developing algorithms for detecting AI-generated text and constructing relevant datasets.However, in the domain of Chinese corpora, challenges remain, including limited model diversity and data homogenization. This paper proposes C-ReD: a comprehensive Chinese benchmark for AI-generated text detection derived from real-world prompts. Experiments show that C-ReD not only supports reliable in-domain detection, but also demonstrates strong generalization to unseen LLMs and external Chinese datasets.
Original Abstract (excerpt)
> Recently, large language models (LLMs) are capable of generating highly fluent textual content. While they offer significant convenience to humans, they also introduce various risks, like phishing and academic dishonesty. Numerous research efforts have been dedicated to developing algorithms for detecting AI-generated text and constructing relevant datasets.--- *Auto-collected on 2026-04-15*