论文概要
研究领域: ML
作者: Junjie Zhang, Hui Liu, Kecheng Chen
发布时间: 2026-08-28
arXiv: 2508.11367
中文摘要
基于LLM的智能体正越来越多地部署在产品级执行环境中,在此环境下,越狱攻击可能触发有害的工具使用和持久状态变更,造成比不安全文本生成更大的风险。现有自动红队方法通常依赖固定攻击,而近期的智能体攻击者协调多个越狱工具,通过基于轨迹的检索展现出更强的潜力。然而,这种检索可能因检索偏差和工具贡献不清而重用误导性经验,且完整轨迹增加了上下文开销并降低了可解释性。我们提出RedEvoAgent——一种黑盒红队智能体,将跨案例攻击轨迹蒸馏为简洁、人类可读的攻击技能。攻击技能通过工具有效性分析和用于技能更新的决策工具归因来自适应演化,并通过验证棘轮仅保留提升验证性能的更新。在多个基准、目标模型和目标执行环境上的实验表明,RedEvoAgent优于固定和智能体基线,提升工具效率,并能跨攻击者模型和目标执行环境迁移。
原文摘要
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text生成 alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval. However, such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability. We propose RedEvoAgent, a black-box red-teaming agent that distills cross-case attack trajectories into a concise, human-readable attack skill. The attack skill adaptively evolves through tool-effectiveness profilin...
自动采集于 2026-08-29
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。