[论文] How Language Models Organize and Structure Moral Knowledge
研究领域: NLP 作者: Orion Reblitz-Richardson 发布时间: 2026-08-27 arXiv: 2608.27402
论文概要
研究领域: NLP 作者: Orion Reblitz-Richardson 发布时间: 2026-08-27 arXiv: 2608.27402
中文摘要
大语言模型如何组织道德知识?模型能广泛检测道德内容,但检测只是低门槛。我们询问它们是否更进一步,区分不同的道德基础并在几何上组织它们之间的关系。我们在开放权重语言模型上训练六个独立线性探针,每个对应道德基础理论(MFT)的一个类别(关怀/伤害、公平/欺骗、自由/压迫、忠诚/背叛、权威/颠覆、圣洁/堕落),并检查结果方向在表示空间中如何相互关联。我们发现这些方向既不坍缩成单一道德检测器,也不相互隔离。相反,它们跨越接近最大数量的独立维度,同时共享一个正向共同成分。共同成分是整合的特征,相对于匹配的非道德概念电池(平均成对余弦0.26 vs 0.013)它是道德特定的。该几何结构跨架构和规模一致,并在预训练早期达到整合状态。
原文摘要
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of ...
*自动采集于 2026-08-30*
#论文 #arXiv #NLP #小凯