Summary
CoALFake is a novel approach for cross-domain fake news detection proposed by Esma Aimeur, Gilles Brassard, and Dorsaf Sallami. It combines human-LLM co-annotation with domain-aware active learning to overcome two major limitations of existing systems: heavy reliance on costly labelled data and information loss from rigid domain categorization. Large language models provide scalable, low-cost annotation while human oversight ensures label reliability. Domain embedding techniques dynamically capture both domain-specific nuances and cross-domain patterns, enabling a domain-agnostic model, and a domain-aware sampling strategy prioritizes diverse domain coverage when selecting samples. Experiments across multiple datasets show CoALFake consistently outperforms various baselines, even with minimal human supervision, highlighting human-LLM co-annotation as a highly cost-effective strategy for generalizable fake news detection.
Overview
- Field: Machine Learning
- Authors: Esma Aimeur, Gilles Brassard, Dorsaf Sallami
Introduction
The proliferation of fake news across diverse domains highlights critical limitations in current detection systems, which often exhibit narrow domain specificity and poor generalization. Existing cross-domain approaches face two key challenges:1. Reliance on labelled data, which is frequently unavailable and resource-intensive to acquire.
2. Information loss caused by rigid domain categorization or neglect of domain-specific features.
Method
To address these issues, the authors propose CoALFake, a novel approach for cross-domain fake news detection that integrates Human-Large Language Model (LLM) co-annotation with domain-aware Active Learning (AL):
- LLM-assisted annotation: LLMs are employed for scalable, low-cost annotation while maintaining human oversight to ensure label reliability.
- Domain embeddings: Domain embedding techniques dynamically capture both domain-specific nuances and cross-domain patterns, enabling the training of a domain-agnostic model.
- Domain-aware sampling: A sampling strategy optimizes sample acquisition by prioritizing diverse domain coverage.
Results
Experimental results across multiple datasets demonstrate that the proposed approach consistently outperforms various baselines, even with minimal human oversight. The results emphasize that human-LLM co-annotation is a highly cost-effective approach that delivers excellent performance.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177169614