[论文] How Language Models Organize and Structure Moral Knowledge

研究领域: NLP 作者: Orion Reblitz-Richardson 发布时间: 2026-08-27 arXiv: 2608.27402

论文概要

研究领域: NLP 作者: Orion Reblitz-Richardson 发布时间: 2026-08-27 arXiv: 2608.27402

中文摘要

大语言模型如何组织道德知识?模型能广泛检测道德内容,但检测只是低门槛。我们询问它们是否更进一步,区分不同的道德基础并在几何上组织它们之间的关系。我们在开放权重语言模型上训练六个独立线性探针,每个对应道德基础理论(MFT)的一个类别(关怀/伤害、公平/欺骗、自由/压迫、忠诚/背叛、权威/颠覆、圣洁/堕落),并检查结果方向在表示空间中如何相互关联。我们发现这些方向既不坍缩成单一道德检测器,也不相互隔离。相反,它们跨越接近最大数量的独立维度,同时共享一个正向共同成分。共同成分是整合的特征,相对于匹配的非道德概念电池(平均成对余弦0.26 vs 0.013)它是道德特定的。该几何结构跨架构和规模一致,并在预训练早期达到整合状态。

原文摘要

How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of ...


*自动采集于 2026-08-30*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens