Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge
研究领域: ML 作者: Yuanchen Bai, Zijian Ding, Angelique Taylor 发布时间: 2026-09-11 arXiv: 2509.05822
论文概要
研究领域: ML 作者: Yuanchen Bai, Zijian Ding, Angelique Taylor 发布时间: 2026-09-11 arXiv: 2509.05822
中文摘要
生成式AI智能体的持续部署不仅要求孤立任务成功。智能体必须在重复交互、变化条件和共享工作流中对人依赖的情况下保持有用性,尤其是在技术、人力和运营中断随时间累积时。我们提出操作韧性和体贴参与作为评估此类智能体的两个互补方面:前者捕捉智能体如何从阻塞工作中恢复,同时保留进度并沟通其限制;后者捕捉其适应如何顾及受影响的人、角色边界和周围工作流。然而在累积挑战下两者均研究不足。我们研究120条模拟医疗轨迹,跨两个生成式AI模型和12个利益相关者派生任务,在轻、中、重挑战下进行比较。我们比较文本行动计划、提示内部评估和定量结构化工作量与情感报告,以检验挑战累积时智能体行为和报告状态如何变化。关于操作韧性,智能体从自主恢复转向更大程度依赖人,同时在结构化报告中报告增加的工作量和负面情感,但很少在文本回应中表达压力。关于体贴参与,智能体从任务聚焦适应扩展到任务重构、关注他人、角色边界调整和更广泛协调,在行动和内部评估中呈现不同模式。基于这些发现,我们推导出五个部署困境,涉及坚持、注意力、角色边界、状态披露和升级,需要利益相关者明确规范,进一步为学习、情境评估和具身适应提供技术启示。
原文摘要
Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge. We study 120 simulated healthcare trajectories across two generative AI models and tw...
*自动采集于 2026-09-12*
#论文 #arXiv #ML #小凯