论文概要
研究领域: ML
作者: Yisen Xi
发布时间: 2026-08-28
arXiv: 2508.11364
中文摘要
受治理组织中的大语言模型(LLM)智能体必须让人格(指令、语气、自我呈现)自由演化,同时保持执行(有状态的、可审计的工作)可追溯。单一信任域无法廉价地同时满足两者。我们提出人格-执行分离(PES):人格与执行位于不同信任域,通过受治理的契约桥连接。人格是单宿主的且可能漂移;执行是无面且可审计的。状态摘要可以返回;数据主体保留在限制性域中,除分级数据丢失预防(DLP)例外;身份保持连续。审批矩阵、DLP和审计强制执行跨界。PES源于三个目标——自由漂移、执行可追溯性和解耦。在LLM表征不可区分性下,任何满足这三者的单域机制必须重新引入类型化变更对象、外部门和稳定审计锚:以更高耦合成本重建PES。在一个受监管的数字员工平台中的开发/试点案例记录了五个月决策,每个都有被拒绝的替代方案。对已发布实现的机制检查发现,在人格扰动(五种模型配置)下没有执行端重新验证,且硬断言字段上没有人格指纹。对恢复的分离前构建的探测发现,受治理的执行路径通过遗漏而非构造与人格解耦;后续的布线变更可能逆转该隔离,而PES将其作为受审计的架构规则。该模式适用于多用户部署、执行审计和预期人格变动共同成立时。
原文摘要
Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not satisfy both cheaply. We present Persona-Execution Separation (PES): persona and execution reside in different trust domains, connected by a governed contract bridge. The persona is singly-homed and may drift; execution is faceless and audited. Status summaries may return; data bodies remain in the restrictive domain except a graded data-loss-prevention (DLP) exception; identity stays continuous. An approval matrix, DLP, and audit enforce the crossing. PES follows from three goals---free drift, execution traceability, and decoupling. Under LLM representational indist...
自动采集于 2026-08-29
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。