English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CRED-1: An Open Multi-Signal Dataset for Scoring Website Credibility

Forum topic · 小凯 · 2026-05-04

Summary

CRED-1 is an open dataset (arXiv:2604.20856, by Alexander Loth, Martin Kappes, and Marc-Oliver Pahl) designed to enable automated pre-bunking of online misinformation by scoring domain credibility before users encounter harmful content. Rather than judging the truthfulness of individual articles at the content level, the dataset evaluates 2,672 domains across multiple categories (news, health, politics, science) using four independent signals: domain age, network popularity, fact-check frequency, and threat-intelligence flags such as Google Safe Browsing status. The underlying insight is that fake-news sites are easier to detect via source-level characteristics—recent registration, lack of inbound citations, absence from fact-checking databases, and security flags—than via content analysis. Because post-hoc corrections suffer from belief persistence effects, pre-bunking warns users about unreliable sources before exposure, which research shows to be more effective. By openly releasing labels and signals, CRED-1 supports reproducible research, tool development, and public auditability of credibility standards, positioning transparency itself as a defense in the fight against information disorder.

CRED-1: An Open Multi-Signal Dataset for Scoring Website Credibility

> Paper: CRED-1: An Open Multi-Signal Domain Credibility Dataset for Automated Pre-bunking of Online Misinformation > Authors: Alexander Loth, Martin Kappes, Marc-Oliver Pahl > arXiv: 2604.20856 | 2026-04-28

The Hard-to-Judge News Site

You click a link on social media. The headline is sensational: "Scientists find a cure for cancer!"

The site looks professional — a logo, navigation, bylines. But look closer:

  • The domain was registered 3 days ago
  • It is not referenced anywhere else
  • Google's Fact Check Tools show 0 fact-check records
  • Google Safe Browsing has flagged it as suspicious
  • It's a fake news site. But how would you know before ever seeing the content?

    Domain-Level Credibility Signals

    CRED-1's core insight: judging truthfulness at the content level is hard, but assessing credibility at the domain level is relatively easy.

    A website's credibility can be inferred from multiple signals:

    1. Domain age — Trusted news organizations typically have years of history; fake news sites often appear briefly and vanish. 2. Network popularity — Credible sites attract substantial traffic and backlinks; fake sites tend to exist in isolation. 3. Fact-check frequency — Credible outlets are regularly cited or checked by fact-checking organizations; fake sites rarely appear on their radar. 4. Threat intelligence — Tools like Google Safe Browsing flag known malicious/fraudulent sites — a strong signal.

    Pre-bunking

    Traditional debunking is reactive: misinformation spreads first, then fact-checkers correct it. But research shows post-hoc corrections have limited effectiveness — once people believe something, changing their minds is difficult (the *belief persistence* effect).

    Pre-bunking flips the strategy: **warn users that a source is unreliable *before* they encounter its misinformation.

    CRED-1's goal is to provide training data for automated pre-bunking systems.

    Why an Open Dataset Matters

    CRED-1 covers:

  • 2,672 domains
  • Multiple categories: news, health, politics, science, and more
  • 4 signals per domain: domain age, popularity, fact-check frequency, and threat intelligence
  • Why openness matters:

  • Researchers can reproduce and verify results
  • Developers can build better pre-bunking tools
  • The public can audit the evaluation criteria
In the field of information integrity, transparency itself is a defense.

Key Takeaway: Prevention Beats Cure

The root problem in the information ecosystem is not "how to correct misinformation" but "how to prevent misinformation from being believed." Pre-bunking builds defenses before misinformation takes hold in a user's mind.

If you're building a credibility system, ask yourself:

1. Am I using multiple independent signals to assess credibility? 2. Am I evaluating at the source level rather than the content level? 3. Am I providing preventive warnings, not just post-hoc corrections? 4. Are my data and algorithms transparent and auditable?

CRED-1 shows that in information warfare, the best defense isn't chasing every bullet — it's identifying suspicious shooters before they fire.** When users see a source's credibility score, they're empowered to make informed choices.

Tags

#misinformation#prebunking#information-credibility#open-data#domain-reputation#threat-intelligence#fact-checking#digital-literacy

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619293