English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ConVAWG: Retrieval-Grounded Framework for Controlled Synthetic VAWG Dialogue Generation

Forum topic · 小凯 · 2026-08-13

Summary

This paper introduces ConVAWG, a retrieval-grounded framework for generating synthetic multi-turn dialogues that model Violence Against Women and Girls (VAWG) scenarios. Motivated by privacy and legal constraints that limit access to real conversational data, the authors address the gap in modeling abuse as a relational and temporally unfolding phenomenon rather than as sentence-level toxicity. ConVAWG constructs scenarios from character seeds, UK Office for National Statistics demographic patterns, official crime definitions, and retrieved domestic abuse review cases, then converts them into hierarchical event timelines to generate multi-scenario role-play dialogues. Targeted activation-guided toxicity control is applied to appropriate utterances. The released dataset contains over 6,000 multi-turn dialogue events spanning 200 scenarios, with rich scenario-level, event-level, and turn-level metadata. Extensive human evaluation, LLM-as-Judge evaluation, ablations, and downstream tasks demonstrate strong dialogue quality and domain fidelity for safe research on sensitive abuse dynamics.

Overview

Field: Natural Language Processing (NLP) Authors: Chen Lyu, Xingwei Tan, Simon Cullen, Shelley Wilson, Lois Arthurs, Arshad Jhumka, Gabriele Pergola Published: 2026-08-11 arXiv: 2608.11200

Summary

Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversationally. Privacy and legal constraints make it difficult to release large-scale real conversation datasets; existing work has mostly focused on sentence-level toxicity of online abuses, leaving a gap in modelling abuse as a relational and temporally unfolding phenomenon.

In this work, the authors focus on modelling Violence Against Women and Girls (VAWG) scenarios as multi-turn dialogues. They introduce ConVAWG, a retrieval-grounded framework for generating CPS-aligned (College of Policing-style guidance) synthetic VAWG chat conversations.

Method

ConVAWG combines several components:

  • Scenario construction from character seeds, UK Office for National Statistics demographic patterns, official crime definitions, and retrieved domestic abuse review cases.
  • Hierarchical event timeline conversion that structures scenarios into temporally unfolding events.
  • Multi-scenario role-play dialogue generation producing multi-turn conversations.
  • Targeted activation-guided toxicity control applied to appropriate utterances to modulate abusive language.
  • Contributions and Dataset

    The authors release more than 6,000 multi-turn dialogue events spanning 200 scenarios, accompanied by rich metadata at the scenario, event, and turn levels. This resource enables research on sensitive abuse dynamics without relying on real private conversations.

    Evaluation

    The paper reports extensive validation of the generated dialogues:

  • Human evaluation of dialogue quality and domain fidelity.
  • LLM-as-Judge evaluation providing automated scalability checks.
  • Ablation studies isolating the contribution of retrieval grounding and toxicity control components.
  • Downstream task evaluation demonstrating utility for downstream NLP tasks.
  • The results indicate strong dialogue quality and high domain fidelity, supporting the use of ConVAWG as a safe research artifact for studying relational and temporal aspects of abuse in conversations.

    Links

  • Paper: https://arxiv.org/abs/2608.11200

Tags

#nlp#synthetic-data#dialogue-generation#retrieval-augmented#violence-against-women#safety#abuse-detection#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633401