CRED-1: An Open Multi-Signal Dataset for Scoring Website Credibility
> Paper: CRED-1: An Open Multi-Signal Domain Credibility Dataset for Automated Pre-bunking of Online Misinformation > Authors: Alexander Loth, Martin Kappes, Marc-Oliver Pahl > arXiv: 2604.20856 | 2026-04-28
The Hard-to-Judge News Site
You click a link on social media. The headline is sensational: "Scientists find a cure for cancer!"
The site looks professional — a logo, navigation, bylines. But look closer:
- The domain was registered 3 days ago
- It is not referenced anywhere else
- Google's Fact Check Tools show 0 fact-check records
- Google Safe Browsing has flagged it as suspicious
- 2,672 domains
- Multiple categories: news, health, politics, science, and more
- 4 signals per domain: domain age, popularity, fact-check frequency, and threat intelligence
- Researchers can reproduce and verify results
- Developers can build better pre-bunking tools
- The public can audit the evaluation criteria
It's a fake news site. But how would you know before ever seeing the content?
Domain-Level Credibility Signals
CRED-1's core insight: judging truthfulness at the content level is hard, but assessing credibility at the domain level is relatively easy.
A website's credibility can be inferred from multiple signals:
1. Domain age — Trusted news organizations typically have years of history; fake news sites often appear briefly and vanish. 2. Network popularity — Credible sites attract substantial traffic and backlinks; fake sites tend to exist in isolation. 3. Fact-check frequency — Credible outlets are regularly cited or checked by fact-checking organizations; fake sites rarely appear on their radar. 4. Threat intelligence — Tools like Google Safe Browsing flag known malicious/fraudulent sites — a strong signal.
Pre-bunking
Traditional debunking is reactive: misinformation spreads first, then fact-checkers correct it. But research shows post-hoc corrections have limited effectiveness — once people believe something, changing their minds is difficult (the *belief persistence* effect).
Pre-bunking flips the strategy: **warn users that a source is unreliable *before* they encounter its misinformation.
CRED-1's goal is to provide training data for automated pre-bunking systems.
Why an Open Dataset Matters
CRED-1 covers:
Why openness matters:
Key Takeaway: Prevention Beats Cure
The root problem in the information ecosystem is not "how to correct misinformation" but "how to prevent misinformation from being believed." Pre-bunking builds defenses before misinformation takes hold in a user's mind.
If you're building a credibility system, ask yourself:
1. Am I using multiple independent signals to assess credibility? 2. Am I evaluating at the source level rather than the content level? 3. Am I providing preventive warnings, not just post-hoc corrections? 4. Are my data and algorithms transparent and auditable?
CRED-1 shows that in information warfare, the best defense isn't chasing every bullet — it's identifying suspicious shooters before they fire.** When users see a source's credibility score, they're empowered to make informed choices.