[论文] User Model Extraction via Belief Self-Distillation
研究领域: NLP 作者: Ali Holmov, Yiran Huang, Kirill Bykov, Zeynep Akata 发布时间: 2026-09-25 arXiv: 2609.31603
论文概要
研究领域: NLP 作者: Ali Holmov, Yiran Huang, Kirill Bykov, Zeynep Akata 发布时间: 2026-09-25 arXiv: 2609.31603
中文摘要
大语言模型(LLM)会隐式推断用户的属性并相应调整自身行为,但这些“信念”一直难以检视,也难以进行因果层面的操控。我们提出信念自蒸馏(Belief Self-Distillation, BSD),一个统一的读写框架:通过学习一种紧凑的用户表示,打通线性探针与因果探针——它既可以被解码读出,也可以被写回模型。冻结的 LLM 充当自己的教师,从自然对话中蒸馏信念,无需外部标注。与常规探针不同,BSD 分离出的不仅是激活值中存在的信息,更是一个其因果作用可被直接检验的状态。在多个模型家族上,BSD 能忠实恢复用户信念,并实现了远强于对等隐状态引导(hidden-state steering)的干预效果。关键的是,我们发现模型的拒绝行为不仅取决于请求本身,还取决于它对用户意图的推断:固定请求、只改变这一信念,就会改变拒绝行为。我们还揭示了一个惊人的跨模型规律:独立训练的 LLM 在用户表示上收敛于共享的几何结构。这些结果共同表明,隐式用户模型是可读取、且可因果写入的内部状态,对 AI 安全有直接启示——它决定了模型如何基于“自己在与谁对话”来调节安全决策。
原文摘要
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state ste...
*自动采集于 2026-09-29*
#论文 #arXiv #NLP #小凯