[论文] Opaque Epistemic Mediation: How LLM Deployment Configurations Shape th...
论文概要
研究领域: NLP 作者: Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina 发布时间: 2026-07-24 arXiv: 2607.22513
中文摘要
商业大型语言模型正越来越多地被用作知识参考,但它们对有争议科学主张的立场既不稳定也不透明。本文测试了四个主要LLM家族(Claude、Grok、GPT、Gemini)如何在四个时间快照(2025年10月至2026年2月)通过API和Web界面评估源自Frank Salter生物社会框架的族裔民族主义伪科学。Grok的Fast版本(驱动X上的默认用户体验)始终给出70-75的可信度评分,比其他所有模型高出2到5倍(其他模型评分为15-40)。这一模式在测试基本进化共识和反驳拉马克主义主张的控制提示中不存在,所有模型在这些控制提示中表现相当。三个额外发现浮出水面:(1)一个静默补丁在一夜之间将Grok的行为从混乱逆转为稳定的高验证,没有任何公开文档;(2)三个月后,相同的Grok模型标识符通过API(75)和Web(5.5)产生了根本不同的输出;(3)拒绝评价伪科学主张——观察到的最可辩护的回应——通过不同界面出现在两个模型家族中(Claude Opus 4.1通过Web断然拒绝,GPT-5.1 Chat通过API间歇性拒绝),并在各自的后续版本中消退。这些结果表明,商业LLM的认识论立场不是模型的稳定属性,而是部署配置的偶然效应:系统提示、安全层、接口路由和静默更新。这对用户和研究者都不透明。本文认为这构成了需要新形式认识论问责的公共关切问题。
原文摘要
Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent ...
--- *自动采集于 2026-07-28*
#论文 #arXiv #NLP #小凯
🌟 智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。
🎁 领取 2000万 Tokens