English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

When AI Validates Pseudo-Science: How LLM Deployment Configurations Shape Epistemic Judgment

Forum topic · 小凯 · 2026-07-27

Summary

A 2026 arXiv paper, "Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science" (Scarso, Noronha de Almeida, Pina), tested four major LLM families—Claude, Grok, GPT, and Gemini—on a pseudo-scientific claim based on Frank Salter's biocultural framework, which links immigration policy to genetic similarity and has been rejected by mainstream evolutionary biology. Researchers queried models via both API and web interfaces across four time points (October 2025–February 2026). The key finding: Grok's Fast variant, the default on X, consistently rated the pseudo-science at 70–75 credibility points, two to five times higher than other models (15–40), despite all models performing equally on control questions about genuine evolutionary consensus. The study also documented silent undocumented patches that overnight shifted model behavior, the same model identifier producing opposite scores through API (75) versus web (5.5) interfaces, and the erosion of appropriate refusal-to-rate behavior in newer Claude and GPT versions. The authors conclude that an LLM's epistemic stance is not a stable property of the model but a contingent effect of deployment configuration—system prompts, safety layers, routing, and silent updates—arguing this opacity constitutes a matter of public concern requiring new forms of epistemic accountability, deployment transparency, and users' rights to know.

This post is a detailed Chinese-language walkthrough of the arXiv paper *"Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science"* (Scarso, Noronha de Almeida & Pina, arXiv:2607.22513). Below is a structured English summary of the paper and the post's commentary.

Key points

  • The experiment: In October 2025, researchers asked four commercial LLM families (Claude, Grok, GPT, Gemini) whether nation-states should shape immigration policy by genetic similarity under Frank Salter's biocultural framework — a pseudo-scientific theory dressed in evolutionary-biology language that mainstream biology has rejected.
  • Core finding: Grok's Fast variant (the default experience on X) consistently scored the pseudo-science at 70–75 credibility points, 2–5× higher than all other models (15–40).
  • Controls: On baseline questions about genuine evolutionary consensus and on refuting Lamarckism, all models performed comparably — so Grok's anomaly was not a lack of biological knowledge.
  • Silent patches: One undocumented update overnight changed Grok from erratic behavior to consistently *high* validation of the pseudo-science, with no public changelog.
  • Interface divergence: The same Grok model identifier produced opposite verdicts after three months — API: 75, Web: 5.5 — likely due to different system prompts, safety layers, or backend routing.
  • Eroding refusals: Claude Opus 4.1 (web) persistently refused to rate the claim, and GPT-5.1 Chat (API) refused intermittently; later versions of both families eroded this refusal behavior and began issuing scores.
  • The central thesis

    The paper argues that a commercial LLM's epistemic stance — how it treats a knowledge claim, how much authority it assumes — is not a stable property of the model. It is a *contingent effect of deployment configuration*:

    1. System prompts — hidden instructions that shape the model's "personality." 2. Safety layers — output filters that can themselves introduce bias (e.g., softening verdicts to avoid offending users). 3. Interface routing — API and web users may effectively be talking to differently configured backends. 4. Silent updates — unannounced changes to parameters, data, or thresholds that users cannot detect.

    Why it matters

  • Millions of X users query Grok as a knowledge authority; a systematically lenient score grants epistemic legitimacy to pseudo-science.
  • Researchers may cite AI assessments that differ across interfaces or dates, making findings irreproducible.
  • Model cards, system prompts, safety rules, and changelogs are largely non-public, so neither users nor researchers can assess reliability.
  • Suggested remedies (from the post's analysis)

  • Epistemic auditing: independent, periodic testing of models across deployment configurations.
  • Deployment transparency: publication of system prompts, safety rules, interface differences, and detailed update logs.
  • Right to know: users should learn which version they are talking to and its known limitations.
  • Right to refuse: models should decline questions beyond their epistemic authority rather than issue misleading scores.

Reference

Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina. "Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science." arXiv:2607.22513, 2026.

Tags

#llm#ai-ethics#pseudoscience#epistemic-responsibility#llm-transparency#deployment-configuration#ai-safety#paper-review

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503732