[论文] Necessary or Sufficient? Evaluating LLM Explanations With Behavioural ...

研究领域: ML 作者: Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri 发布时间: 2026-09-04 arXiv: 2609.05385

论文概要

研究领域: ML 作者: Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri 发布时间: 2026-09-04 arXiv: 2609.05385

中文摘要

在智能体工作流中运行的LLM决策组件常产生行动相关建议或判断,并附带解释。操作者可能用这些命名因素监控系统、诊断错误或决定何时升级输出。这类使用假设解释与组件的可观测决策行为一致。本文测试两种解释解读:必要性——改变某因素会改变输出;充分性——保留该因素同时移除其他可变信息会维持输出。我们在两个合成用例中评估:向客户推荐顾问和判断提示的有害性/风险。模型返回输出及影响最大的前三个因素。受控黑盒干预通过测量改变因素导致输出变化频率估计必要性得分,通过测量保留因素维持输出的频率估计充分性得分。跨Claude/GPT/Gemini八款模型,顾问推荐任务中引用排序与必要性/充分性得分的平均Spearman相关分别为0.349和0.354,提示监控任务为0.431和0.580。此外,未引用因素得分超过最低引用因素的比率:顾问必要性57.6%、充分性58.1%,提示监控25.8%和8.9%。引用前三包含有用信息,但在必要性或充分性下不能可靠识别影响最强的三个因素。该框架为智能体监督中使用的解释提供了黑盒可靠性检验。

原文摘要

LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named factors to monitor a system, diagnose errors, or decide when to escalate an output. Such use assumes that the explanations agree with the component's observable decision behaviour. We test two interpretations of the named factors: necessity, meaning that changing a factor would change the output, and sufficiency, meaning that retaining it while removing other changeable information would preserve the output. We evaluate these interpretations in two synthetic use cases: recommending advisors to clients and judging prompts for harmfulness or risk. Models return an output and the top three factors that most influenc...


*自动采集于 2026-09-09*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens