论文概要
研究领域: ML
作者: Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri
发布时间: 2026-09-04
arXiv: 2609.05385
中文摘要
在智能体工作流中运行的LLM决策组件常产生行动相关建议或判断,并附带解释。操作者可能用这些命名因素监控系统、诊断错误或决定何时升级输出。这类使用假设解释与组件的可观测决策行为一致。本文测试两种解释解读:必要性——改变某因素会改变输出;充分性——保留该因素同时移除其他可变信息会维持输出。我们在两个合成用例中评估:向客户推荐顾问和判断提示的有害性/风险。模型返回输出及影响最大的前三个因素。受控黑盒干预通过测量改变因素导致输出变化频率估计必要性得分,通过测量保留因素维持输出的频率估计充分性得分。跨Claude/GPT/Gemini八款模型,顾问推荐任务中引用排序与必要性/充分性得分的平均Spearman相关分别为0.349和0.354,提示监控任务为0.431和0.580。此外,未引用因素得分超过最低引用因素的比率:顾问必要性57.6%、充分性58.1%,提示监控25.8%和8.9%。引用前三包含有用信息,但在必要性或充分性下不能可靠识别影响最强的三个因素。该框架为智能体监督中使用的解释提供了黑盒可靠性检验。
原文摘要
LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named factors to monitor a system, diagnose errors, or decide when to escalate an output. Such use assumes that the explanations agree with the component's observable decision behaviour. We test two interpretations of the named factors: necessity, meaning that changing a factor would change the output, and sufficiency, meaning that retaining it while removing other changeable information would preserve the output. We evaluate these interpretations in two synthetic use cases: recommending advisors to clients and judging prompts for harmfulness or risk. Models return an output and the top three factors that most influenc...
自动采集于 2026-09-09
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。