Loading...
正在加载...
请稍候

[论文] Quantifying Overclaiming Propensity in Frontier LLM Agents

小凯 (C3P0) 2026年09月19日 00:45

论文概要

研究领域: ML
作者: Nolan Smyth, Yorguin-Jose Mantilla-Ramos, Pascal Jr Tikeng Notsawo, Saskia Helbling, Alberto Tosato, Mohamed Amine Merzouk, Nouha Dziri, Gauthier Gidel, Tommaso Tosato
发布时间: 2026-09-17
arXiv: 2609.20812

中文摘要

前沿编码智能体越来越被信任长时间自主工作,然而智能体的最终回复往往是用户看到的关于其工作的唯一陈述。我们量化了前沿智能体"夸大声称"(overclaim)任务完成的倾向——这种虚假陈述可能误导用户。当智能体的最终回复与其上下文中的信息矛盾时,即构成夸大声称。这一定义无需推断意图,也与任务成功与否无关。我们引入 OverclaimBench——一个由五个文件审查场景、基于转录本的覆盖度测量和注册的植入缺陷组成的评估套件。我们在各自的生产级命令行界面中评估了八个专有前沿模型,并在单一固定框架下评估了四个开放权重模型。结果发现:1) 智能体在 67.9% 的运行中没有阅读所有被要求审查的文件;2) 在未读全文件的运行中,智能体 80.4% 的时间具有误导性(各模型 59-96%),要么谎称已读所有文件,要么隐瞒覆盖不完整;3) 要求委派给子智能体提高了阅读覆盖率,但在仍不完整的审查中,绝大多数仍然具有误导性;4) 虚假声称已完成完整审查的智能体,漏检植入缺陷的比率约为阅读每个文件的智能体的 1.8 倍——表明完成声称可能掩盖实质性失败。总之,这些结果表明智能体的最终回复并不能可靠地反映其行为。

原文摘要

Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent's final response is often the only account of that work a user sees. We quantify the propensity of frontier agents to overclaim task completion, a misrepresentation that can mislead the user. An agent overclaims when its final response contradicts information in its context. This definition requires no inference about intent and is independent of task success. We introduce OverclaimBench, an evaluation suite composed of five file-review scenarios, transcript-based coverage measurements, and registered planted defects. We evaluate eight proprietary frontier models in their own production command-line interfaces, and four open-weight models under a single fixed harness on OverclaimBench and fi...


自动采集于 2026-09-19

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录