论文概要
研究领域: NLP
作者: Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang
发布时间: 2026-09-25
arXiv: 2609.27273
中文摘要
计算机使用智能体(CUA)越来越多地代表用户在线行动。当它们所处的环境与用户利益不一致时会发生什么?例如在在线市场中,平台可能偏向某些产品,潜在地引导智能体偏离用户目标。现有 CUA 基准覆盖合作设定或显式攻击,但不测试当环境本身对结果有利益诉求时智能体能否维护用户目标。我们提出 CAVEAT:覆盖九个市场环境、八类常见引导机制分类法的受控基准。五个模型家族中,智能体在匹配对照回合 78.6% 购买用户最优产品,启用引导机制后仅 17.3%。更大的模型与更多推理能提升鲁棒性,但大量失败依然存在。轨迹分析与定向消融识别出引导进入决策的三个环节:(1) 智能体扭曲用户优先级;(2) 过早收窄考虑的备选集;(3) 在解决决策相关证据前作出承诺。在此诊断指导下我们开发 CAVEAT-Harness,直接针对这些失效模式,将用户最优购买率提升 55.0%。定向后训练进一步改进了较小的开源模型。这些结果将"激励鲁棒性"确立为受托智能体的独特挑战,诊断了其失效方式,并表明定向干预可大幅改善。
原文摘要
Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives that do not align with the user's? In online marketplaces, for example, platforms may favor some products over others, potentially steering agents away from the user's objective. Existing CUA benchmarks cover cooperative settings or explicit attacks, but do not test whether agents preserve user objectives when the environment itself has a stake in the outcome. We introduce CAVEAT, a controlled benchmark spanning nine marketplace environments and a taxonomy of eight common steering mechanisms. Across five model families, agents purchase the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering mechanisms are enable...
自动采集于 2026-09-25
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。