Paper Overview
- Field: AI
- Authors: Yiting Huang, Wenting Zhu, Zekun Wang, et al.
- Published: 2026-05-28
- arXiv: 2605.27584
- Multimodal content moderation
- Interpretability
- Algorithmic fairness
- Dual-use risks of generative AI
Abstract
The proliferation of social media platforms and online communities has inadvertently catalyzed the spread of cyberbullying, hate speech, and other forms of online toxicity, making the effective governance of such harm a critical societal and computational challenge. While significant strides have been made in automating content moderation, existing research predominantly treats cyberbullying governance as passive, isolated detection at the post level. This reductionist view overlooks the continuous behavioral dynamics of users, the structural diffusion of toxic events, and the critical need for proactive mitigation. To bridge these gaps, this paper proposes a unified full-lifecycle governance framework that shifts the paradigm of cyberbullying governance from isolated static detection toward integrated, continuous, and proactive moderation.
Framework: Four Interconnected Stages
Drawing on cyberbullying research and adjacent fields, the paper systematically synthesizes recent literature across four stages:
1. Content identification — detecting toxic and bullying content. 2. User and behavior modeling — capturing users' continuous behavioral dynamics beyond single posts. 3. Diffusion dynamics and early warning — understanding the structural spread of toxic events and predicting escalation. 4. Intervention and governance — proactive mitigation strategies rather than reactive removal alone.
Datasets, Evaluation, and Open Challenges
The paper also reviews available datasets and evaluation practices, and discusses emerging challenges, including:
*Auto-collected on 2026-05-29.*