[论文] RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill ...

研究领域: ML 作者: Junjie Zhang, Hui Liu, Kecheng Chen 发布时间: 2026-08-28 arXiv: 2508.11367

论文概要

研究领域: ML 作者: Junjie Zhang, Hui Liu, Kecheng Chen 发布时间: 2026-08-28 arXiv: 2508.11367

中文摘要

基于LLM的智能体正越来越多地部署在产品级执行环境中,在此环境下,越狱攻击可能触发有害的工具使用和持久状态变更,造成比不安全文本生成更大的风险。现有自动红队方法通常依赖固定攻击,而近期的智能体攻击者协调多个越狱工具,通过基于轨迹的检索展现出更强的潜力。然而,这种检索可能因检索偏差和工具贡献不清而重用误导性经验,且完整轨迹增加了上下文开销并降低了可解释性。我们提出RedEvoAgent——一种黑盒红队智能体,将跨案例攻击轨迹蒸馏为简洁、人类可读的攻击技能。攻击技能通过工具有效性分析和用于技能更新的决策工具归因来自适应演化,并通过验证棘轮仅保留提升验证性能的更新。在多个基准、目标模型和目标执行环境上的实验表明,RedEvoAgent优于固定和智能体基线,提升工具效率,并能跨攻击者模型和目标执行环境迁移。

原文摘要

LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text生成 alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval. However, such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability. We propose RedEvoAgent, a black-box red-teaming agent that distills cross-case attack trajectories into a concise, human-readable attack skill. The attack skill adaptively evolves through tool-effectiveness profilin...


*自动采集于 2026-08-29*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens