[论文] Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse

研究领域: ML 作者: Shreya Gopalan, Devansh Singh, Sundaraparipurnan Narayanan 发布时间: 2026-09-25 arXiv: 2609.26836

论文概要

研究领域: ML 作者: Shreya Gopalan, Devansh Singh, Sundaraparipurnan Narayanan 发布时间: 2026-09-25 arXiv: 2609.26836

中文摘要

智能体 AI 系统正越来越多地采用集成多种工具的自动化流水线。虽然先前的研究与基准主要关注任务成功率与完成度,但针对智能体与工具交互本身的研究——尤其在生物学科智能体工作流中——仍较有限。本研究调查一类特定故障:工具调用看似成功,但通过 API/封装层返回的部分或全部信息或功能不完整甚至缺失,且用户或智能体未收到任何关于信息缺失的通知。我们称之为"静默故障"——用户与智能体都不知道故障已经发生。为此我们开发了一种审计机制,通过检查 ToolUniverse 环境中集成的 15 个科学工具(及其 API 文档和工具文档)来识别此类故障(ToolUniverse 仅作实验环境,而非研究对象)。研究围绕 7 个故障位置(failure locus)展开,刻画故障在链路中发生的环节。我们观察到 91 个故障(经 LLM 候选发现与自动化测试后人工验证),最常见的是数据或字段缺失,以及搜索、过滤或排序标准不一致。91 个故障中多数发生在 API 层(51 个)或封装层(25 个),并可能在下游放大为静默故障。结果表明:静默故障起源于上游事件,向下游传播成表面上有效的科学输出。我们提出"上下文可靠性"(contextual reliability)概念来应对此类故障,并建议在智能体-工具交互全链路中开展测试、披露、监控与度量。

原文摘要

Agentic AI systems are increasingly adopting automated pipelines that integrate multiple tools. While prior research and benchmarks have studied about task success and task completion of these agentic systems, the research about agent to tool interaction, specifically in biology agentic workflow is limited. This study investigates specific failures in agent to tool interaction where a tool invocation appears successful, some or all of the information or functionality from the tool via API/ wrapper is incomplete or missing and there are no communications / notifications to the user or the agent about such missing information. We call this a silent failures as the user or the agents are not aware that such failure has occurred. For the purposes of this study we developed an audit mechanism t...


*自动采集于 2026-09-25*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens