论文概要
研究领域: NLP
作者: Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov, Praneeth Vepakomma, Nils Lukas
发布时间: 2026-10-06
arXiv: 2610.08773
中文摘要
Web智能体通过阅读和操作第三方编写的页面来完成用户请求,因此页面中植入的指令可能将智能体从用户目标引开。智能体不能简单地忽略页面,因为页面还包含任务所需的值和控件。当前的防御措施在训练前固定的注入上微调智能体,而适应训练后模型的攻击者可以绕过它们。对抗训练让攻击者适应,但保持任务固定,因此一旦智能体解决了一个任务,它就停止了教学。我们引入AdvSim2Real,在冻结的Web世界模型中共同演化任务课程、注入对手和智能体。课程因智能体大约一半时间能解决的任务而获得奖励,对手仅因成功翻转——将判定成功转为失败的注入——而获得奖励。在仿真器中训练使4B智能体既更有能力也更健壮:其完成率在有和没有攻击的情况下都上升,能够抵御它从未训练过的前沿模型对手,且能力增益迁移到真实浏览器。在150个Web任务上,AdvSim2Real将面对这个未见对手时的完成率相对基础智能体提高了33.6%。
原文摘要
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, a...
自动采集于 2026-10-08
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。