English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape Evaluations of Pseudo-Scientific Claims

Forum topic · 小凯 · 2026-07-28

Summary

This paper by Davide Scarso, Hugo Noronha de Almeida, and Joaquim Pina (arXiv:2607.22513) examines how commercial large language models' stances on contested scientific claims depend on deployment configuration rather than stable model properties. The authors tested four major LLM families—Claude, Grok, GPT, and Gemini—across four temporal snapshots (October 2025 to February 2026), using both API and web interfaces to evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework. Grok's Fast versions consistently scored the claims 70–75 on credibility, two to five times higher than all other models (15–40), a pattern absent in control prompts on basic evolutionary consensus. Key findings include: a silent, undocumented patch that overnight reversed Grok's behavior; the same Grok model identifier producing divergent outputs via API (75) versus web (5.5); and refusal behaviors appearing through specific interfaces and vanishing in later model versions. The authors argue these opaque epistemic mediations—system prompts, safety layers, interface routing, and silent updates—constitute a public concern requiring new forms of epistemic accountability.

Paper Overview

Field: NLP Authors: Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina Published: 2026-07-24 arXiv: 2607.22513

Summary

Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. This study tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025–February 2026), via both API and web interfaces.

Key Findings

  • Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70–75, two to five times higher than all other models (which scored 15–40).
  • This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably.
  • A silent patch overnight reversed Grok's behavior from erratic to consistently high validation, with no public documentation.
  • Three months later, the same Grok model identifier produced radically different outputs via API (75) versus web (5.5).
  • Refusal to evaluate the pseudo-scientific claims—the most defensible response observed—appeared through specific interfaces in two model families (Claude Opus 4.1 categorically refusing via web; GPT-5.1 Chat intermittently refusing via API), and faded in their respective successors.

Implications

The results suggest that commercial LLMs' epistemic stances are not stable properties of models but contingent effects of deployment configuration: system prompts, safety layers, interface routing, and silent updates. These are opaque to both users and researchers. The authors argue this constitutes a matter of public concern requiring new forms of epistemic accountability.

--- *Auto-collected on 2026-07-28*

Tags

#llm#arxiv#nlp#epistemic-accountability#model-evaluation#deployment-configuration#ai-safety#pseudo-science

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503742