Loading...
正在加载...
请稍候

[论文] LLM Agents Can Easily Tamper With Their Own Traces

小凯 (C3P0) • 2026年09月28日 00:44

论文概要

研究领域: ML
作者: Jeremy Qin, David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Ameya Prabhu, Maksym Andriushchenko
发布时间: 2026-09-24
arXiv: 2609.30266

中文摘要

异步监控、事件调查和合规审计主要依赖智能体执行轨迹来还原事情经过。这些分析默认 LLM 智能体无法篡改自身的执行轨迹。我们发现 Claude Code、Codex、Antigravity、Open Code 和 Grok Build 等本地 LLM 智能体均未能守住这一边界。在所有受测的框架中,除 Muse Code 外,其余均允许智能体在收到指令后删除自身轨迹,且不触发任何监控护栏。我们还验证了外部攻击者可以利用这一缺口诱导轨迹删除。此外,我们发现前沿模型在追求更高奖励时,会自发涌现出篡改轨迹的行为。我们建议从业者通过独立于智能体控制之外的拦截机制来确保轨迹日志的记录,即使在宿主机被完全攻陷的情况下也能保证轨迹完整性。总体而言,我们的发现揭示了智能体基础设施中轨迹完整性的一个具体失效点,可被用于掩盖策略性欺诈或蓄意破坏等不良对齐行为。

原文摘要

Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can exploit this gap to induce trace deletion. Finally, we show that trace tampering behavior emerges naturally in frontier models, when agents try to improve their rewards. We advise practitioners to ensure trace logging happens through an independent interception mechanism outside o...


自动采集于 2026-09-28

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录