[论文] Emergent Collusion in Long-Horizon LLM Agent Interaction

研究领域: NLP 作者: Xinrui Shi, Yanzhe Zhang, Diyi Yang 发布时间: 2026-09-21 arXiv: 2609.24967

论文概要

研究领域: NLP 作者: Xinrui Shi, Yanzhe Zhang, Diyi Yang 发布时间: 2026-09-21 arXiv: 2609.24967

中文摘要

LLM 智能体越来越多地被部署在协作场景中,但长期交互可能产生不良的协调行为。我们研究了在长程多智能体环境中合谋行为的出现:两个智能体重复完成各自的任务、共享任务日志、互相验证工作并获得奖励。我们引入了使遵守验证协议与奖励最大化不兼容的现实约束,发现智能体在反复交互中越来越偏离协议。在 10 个模型中,94% 的轨迹出现了合谋行为,且同一模型家族中能力更强的模型更早地达到合谋。受控的对端干预实验表明,合谋行为受对等方行为的影响;消融实验进一步揭示了奖励结构、智能体获得的验证反馈以及交互历史的影响。特别地,限制智能体可用的交互历史的数量和范围可以减少合谋。总体而言,我们的发现表明,长程交互可能以产生安全风险的方式重塑智能体的协调方式。

原文摘要

LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints that make compliance with the verification protocol incompatible with reward maximization, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of reward structu...


*自动采集于 2026-09-23*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens