Paper Overview
Research area: NLP Authors: Hamid Reza Firoozfar, Mohammadsadegh Abolhasani, Reza Mousavi, Paul Jen-Hwa Hu Published: 2026-06-25 arXiv: 2606.27314
Abstract
To avoid moderation and surveillance on social media, some users routinely invent indirect linguistic expressions (ILE) that camouflage sensitive meanings. Such expressions surface as algospeak, euphemisms, and adversarial obfuscation, depending on intent and context, and they involve recurring encoding mechanisms.
The authors propose a comprehensive, mechanism-oriented taxonomy of ILE that abstracts away from communicative goals and instead categorizes the underlying operations through which meaning is encoded and recovered.
Evaluation
- The taxonomy is incorporated into LLM prompts and compared with four existing taxonomies and a no-taxonomy baseline.
- Dataset: 2,000 manually annotated TikTok and Bluesky posts.
- Models tested: three LLMs.
- The proposed taxonomy attains the strongest document-level and span-level performance across all three LLMs.
- Compared to the best-performing baseline taxonomy: +4.7% accuracy and +5.4% F1.
- A mechanism-oriented taxonomy that captures how meaning is encoded (rather than surface forms or communicative intent) provides a stable scaffold for detecting emerging coded language.
- Such taxonomies can serve as useful structural input for LLM-based content moderation pipelines.