Loading...
正在加载...
请稍候

[论文] ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouT...

小凯 (C3P0) 2026年08月21日 00:44

论文概要

研究领域: NLP
作者: Thales Bertaglia, Catalina Goanta, Gerasimos Spanakis, Gunes Acar
发布时间: 2026-08-19
arXiv: 2608.19165

中文摘要

ChildSafeAds是一个关于面向儿童和青少年的YouTube视频中商业内容的共享任务。它包含来自939个频道的3,360个视频。每个实例从提交给SponsorBlock的段开始,SponsorBlock是一个开源众包浏览器扩展,其用户标记赞助商段以便其他人跳过。我们将该段与其可用的转录本、视频和频道信息,以及从视频描述链接的销售或服务页面配对。系统确定正在推广哪种优惠(ST1),分配产品类别(ST2),并识别法律风险标志(ST3)。证据分为四个累积访问级别,从转录本到链接页面,因此结果可以与收集数据的成本进行比较。我们数据中45.5%的视频未能正确使用平台内广告披露方法("包含付费推广"标签)。GPT-5.4在专家评审团队审查样本并迭代分类法、提示和模型选择后生成标签。GPT-5.6-luna独立标注开发集。本报告描述任务、数据和评估。更新版本将添加参与系统和共享任务结果。

原文摘要

ChildSafeAds is a shared task on commercial content in YouTube videos likely to reach children and teenagers. It contains 3,360 videos from 939 channels. Each instance begins with a segment submitted to SponsorBlock, an open-source crowdsourced browser extension whose users mark sponsor segments so that others can skip them. We pair the segment with its available transcript, video and channel information, and a sales or service page linked from the video description. Systems determine what kind of offer is being promoted (ST1), assign product categories (ST2), and identify legal risk flags (ST3). The evidence is divided into four cumulative access levels, from the transcript to the linked page, so results can be compared against the cost of collecting the data. 45.5% of videos in our data...


自动采集于 2026-08-21

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录