[论文] Vulnerability Localization Benchmark: Measuring Agentic Security Analy...

研究领域: ML 作者: Aman Priyanshu, Supriti Vijay, Kimia Majd, Xuhong He, Fraser Burch, Takahiro Matsumoto, Jianliang He, Baturay Saglam, Arthur Goldblatt, Zhuoran Ya…

论文概要

研究领域: ML 作者: Aman Priyanshu, Supriti Vijay, Kimia Majd, Xuhong He, Fraser Burch, Takahiro Matsumoto, Jianliang He, Baturay Saglam, Arthur Goldblatt, Zhuoran Yang, Amin Karbasi 发布时间: 2026-09-14 arXiv: 2609.15939

中文摘要

语言模型智能体越来越多地在完整的软件代码库上运行,然而网络安全评估主要衡量它们能否检测、复现或修复漏洞,而不是能否定位相关代码。我们研究漏洞定位:给定一个弱点类别和一个不熟悉的代码库,识别与该弱点相关的实现文件。我们引入漏洞定位基准(VLoc Bench),包含来自290个代码库、六个包生态系统和147个CWE类别的500个真实漏洞。每个任务将安全修复前后的代码库快照配对。在易受攻击的快照上,智能体仅接收CWE描述和只读终端访问权限,必须返回受影响的文件;在已修补的快照上,必须判断记录的漏洞是否已不存在。我们在通用智能体接口下评估了27个语言模型和4个静态分析工具。代码库规模的漏洞定位仍然困难:最强系统的文件F1仅为0.229,38.4%的任务未获得任何被评估模型的正确定位。我们还发现,更强的定位能力并不意味着修复后的可靠行为:能有效识别漏洞文件的系统在已修补的代码库上仍可能报告无根据的位置。这些结果确立了漏洞定位作为一种独特的代码库级能力,为研究安全智能体如何搜索脆弱代码以及何时应克制不报提供了场景。

原文摘要

Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can detect, reproduce, or repair vulnerabilities rather than whether they can locate the relevant code. We study vulnerability localization: given a weakness class and an unfamiliar repository, identify the implementation files associated with that weakness. We introduce the Vulnerability Localization Benchmark (VLoc Bench), comprising 500 real world vulnerabilities from 290 repositories across six package ecosystems and 147 CWE categories. Each task pairs repository snapshots immediately before and after a security fix. On the vulnerable snapshot, an agent receives only the CWE description and read-only terminal access and must return the affected fi...


*自动采集于 2026-09-16*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens