Paper Overview
Field: NLP Authors: Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina Published: 2026-07-24 arXiv: 2607.22513
Summary
Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. This study tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025–February 2026), via both API and web interfaces.
Key Findings
- Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70–75, two to five times higher than all other models (which scored 15–40).
- This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably.
- A silent patch overnight reversed Grok's behavior from erratic to consistently high validation, with no public documentation.
- Three months later, the same Grok model identifier produced radically different outputs via API (75) versus web (5.5).
- Refusal to evaluate the pseudo-scientific claims—the most defensible response observed—appeared through specific interfaces in two model families (Claude Opus 4.1 categorically refusing via web; GPT-5.1 Chat intermittently refusing via API), and faded in their respective successors.
Implications
The results suggest that commercial LLMs' epistemic stances are not stable properties of models but contingent effects of deployment configuration: system prompts, safety layers, interface routing, and silent updates. These are opaque to both users and researchers. The authors argue this constitutes a matter of public concern requiring new forms of epistemic accountability.
--- *Auto-collected on 2026-07-28*