English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CoALFake: Collaborative Active Learning with Human-LLM Co-Annotation for Cross-Domain Fake News Detection

Forum topic · 小凯 · 2026-04-07

Summary

CoALFake is a novel approach for cross-domain fake news detection proposed by Esma Aïmeur, Gilles Brassard, and Dorsaf Sallami. It addresses two key limitations of existing cross-domain methods: reliance on scarce, resource-intensive labelled data, and information loss caused by rigid domain categorization or neglect of domain-specific features. The method combines human-LLM co-annotation, using large language models for scalable, low-cost labeling with human oversight to ensure label reliability, together with domain-aware active learning. Domain embedding techniques dynamically capture domain-specific nuances and cross-domain patterns, enabling training of a domain-agnostic model, while a domain-aware sampling strategy prioritizes diverse domain coverage when acquiring samples. Experiments across multiple datasets show CoALFake consistently outperforms existing baselines even with minimal human supervision, demonstrating that human-LLM co-annotation is a highly cost-effective strategy for building generalizable fake news detection systems.

Paper Overview

Research Area: ML Authors: Esma Aïmeur, Gilles Brassard, Dorsaf Sallami

Abstract

The proliferation of fake news across diverse domains highlights critical limitations in current detection systems, which often exhibit narrow domain specificity and poor generalization. Existing cross-domain approaches face two key challenges: (1) reliance on labelled data, which is frequently unavailable and resource intensive to acquire and (2) information loss caused by rigid domain categorization or neglect of domain-specific features. To address these issues, we propose CoALFake, a novel approach for cross-domain fake news detection that integrates Human-Large Language Model (LLM) co-annotation with domain-aware Active Learning (AL).

Our method employs LLMs for scalable, low-cost annotation while maintaining human oversight to ensure label reliability. By integrating domain embedding techniques, CoALFake dynamically captures both domain-specific nuances and cross-domain patterns, enabling the training of a domain-agnostic model. Furthermore, a domain-aware sampling strategy optimizes sample acquisition by prioritizing diverse domain coverage.

Experimental results across multiple datasets demonstrate that the proposed approach consistently outperforms various baselines. Our results emphasize that human-LLM co-annotation is a highly cost-effective approach that delivers excellent performance, even with minimal human oversight.

Key Contributions

  • Human-LLM co-annotation pipeline: LLMs handle scalable, low-cost labeling while humans provide oversight for label reliability.
  • Domain embeddings: Dynamic capture of both domain-specific nuances and cross-domain patterns for a domain-agnostic model.
  • Domain-aware active learning sampling: Prioritizes diverse domain coverage when selecting samples for annotation.
  • Strong empirical results: Consistently outperforms a range of existing baselines across multiple datasets, even with minimal human supervision.
--- *Auto-collected on 2026-04-07*

Tags

#fake-news-detection#active-learning#llm#cross-domain#human-ai-collaboration#data-annotation#machine-learning#domain-adaptation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169635