English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ChildSafeAds Shared Task 2026: Detecting Commercial Content in Child-Facing YouTube Videos

Forum topic · 小凯 · 2026-08-21

Summary

ChildSafeAds is a shared task (arXiv:2608.19165) focused on detecting commercial content in YouTube videos likely to reach children and teenagers. The dataset contains 3,360 videos from 939 channels, with instances seeded by sponsor segments submitted to SponsorBlock, an open-source crowdsourced browser extension. Each segment is paired with its transcript, video and channel metadata, and a sales or service page linked from the video description. Participating systems perform three subtasks: identifying the type of offer being promoted (ST1), assigning product categories (ST2), and flagging legal risks (ST3). Evidence is organized into four cumulative access levels, from transcripts to linked pages, allowing performance to be weighed against data collection costs. Notably, 45.5% of videos in the dataset fail to correctly use YouTube's in-platform paid-promotion disclosure. Labels were generated by GPT-5.4 after expert review and iterative refinement of taxonomy, prompts, and model selection, with GPT-5.6-luna independently annotating the development set. This report describes the task, data, and evaluation; future versions will include participating systems and results.

Paper Overview

  • Field: NLP
  • Authors: Thales Bertaglia, Catalina Goanta, Gerasimos Spanakis, Gunes Acar
  • Published: 2026-08-19
  • arXiv: 2608.19165
  • Abstract

    ChildSafeAds is a shared task on commercial content in YouTube videos likely to reach children and teenagers. It contains 3,360 videos from 939 channels. Each instance begins with a segment submitted to SponsorBlock, an open-source crowdsourced browser extension whose users mark sponsor segments so that others can skip them. Each segment is paired with its available transcript, video and channel information, and a sales or service page linked from the video description.

    Subtasks

  • ST1: Determine what kind of offer is being promoted
  • ST2: Assign product categories
  • ST3: Identify legal risk flags
  • The evidence is divided into four cumulative access levels, from the transcript to the linked page, so results can be compared against the cost of collecting the data.

    Key Findings

  • 45.5% of videos in the dataset fail to correctly use the in-platform advertising disclosure method (the "includes paid promotion" label).
  • Labels were generated by GPT-5.4 after an expert review team examined samples and iterated on the taxonomy, prompts, and model selection.
  • GPT-5.6-luna independently annotated the development set.
This report describes the task, data, and evaluation. Updated versions will add participating systems and shared task results.

---

*Auto-collected on 2026-08-21.*

Tags

#nlp#shared-task#youtube#child-safety#advertising-disclosure#sponsorblock#dataset#legal-risk

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633740