Loading...
正在加载...
请稍候

[论文] ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialo...

小凯 (C3P0) 2026年08月13日 00:45

论文概要

研究领域: NLP
作者: Chen Lyu, Xingwei Tan, Simon Cullen, Shelley Wilson, Lois Arthurs, Arshad Jhumka, Gabriele Pergola
发布时间: 2026-08-11
arXiv: 2608.11200

中文摘要

合成对话生成为研究敏感领域中的对话动态提供了一种途径,在这些领域真实数据难以获取、发布或标注。潜在的虐待可能发生在 online 或 offline:威胁和胁迫可能直接出现在消息中,而监视、孤立、跟踪和身体暴力等行为可能在对话中被计划、披露或提及。隐私和法律限制使得大规模真实对话数据集的发布变得困难;现有工作主要集中在 online 虐待的句子级毒性上,留下了将虐待建模为一种关系性和时间上展开的现象的空白。在这项工作中,我们专注于将针对妇女和女童的暴力(VAWG)场景建模为多轮对话。我们引入了ConVAWG,一个用于生成CPS对齐的合成VAWG聊天对话的检索基础框架。ConVAWG从角色种子、英国国家统计局报告的人口统计模式、官方犯罪定义和检索到的家庭暴力审查案例中构建场景;将它们转换为分层事件时间线;生成多场景角色扮演对话;并对适当的话语应用目标化的激活引导毒性控制。我们发布了跨越200个场景的6,000多个多轮对话事件,包含丰富的场景级、事件级和轮次级元数据。广泛的人工评估、LLM-as-Judge评估、消融实验和下游任务显示了强大的对话质量和领域保真度。

原文摘要

Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversationally. Privacy and legal constraints make it difficult the release of large-scale real conversation datasets; existing work has mostly focused on sentence-level toxicity of online abuses, leaving a gap in modelling abuse as a relational and temporally unfolding phenomenon. In this work, we focus on modelling Violence Against Women and Girls (VAWG) scenarios as multi-turn dialogues. We introduce Con...


自动采集于 2026-08-13

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录