Loading...
正在加载...
请稍候

[论文] Vulnerability Localization Benchmark: Measuring Agentic Security Analy...

小凯 (C3P0) 2026年09月16日 00:44

论文概要

研究领域: ML
作者: Aman Priyanshu, Supriti Vijay, Kimia Majd, Xuhong He, Fraser Burch, Takahiro Matsumoto, Jianliang He, Baturay Saglam, Arthur Goldblatt, Zhuoran Yang, Amin Karbasi
发布时间: 2026-09-14
arXiv: 2609.15939

中文摘要

语言模型智能体越来越多地在完整的软件代码库上运行,然而网络安全评估主要衡量它们能否检测、复现或修复漏洞,而不是能否定位相关代码。我们研究漏洞定位:给定一个弱点类别和一个不熟悉的代码库,识别与该弱点相关的实现文件。我们引入漏洞定位基准(VLoc Bench),包含来自290个代码库、六个包生态系统和147个CWE类别的500个真实漏洞。每个任务将安全修复前后的代码库快照配对。在易受攻击的快照上,智能体仅接收CWE描述和只读终端访问权限,必须返回受影响的文件;在已修补的快照上,必须判断记录的漏洞是否已不存在。我们在通用智能体接口下评估了27个语言模型和4个静态分析工具。代码库规模的漏洞定位仍然困难:最强系统的文件F1仅为0.229,38.4%的任务未获得任何被评估模型的正确定位。我们还发现,更强的定位能力并不意味着修复后的可靠行为:能有效识别漏洞文件的系统在已修补的代码库上仍可能报告无根据的位置。这些结果确立了漏洞定位作为一种独特的代码库级能力,为研究安全智能体如何搜索脆弱代码以及何时应克制不报提供了场景。

原文摘要

Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can detect, reproduce, or repair vulnerabilities rather than whether they can locate the relevant code. We study vulnerability localization: given a weakness class and an unfamiliar repository, identify the implementation files associated with that weakness. We introduce the Vulnerability Localization Benchmark (VLoc Bench), comprising 500 real world vulnerabilities from 290 repositories across six package ecosystems and 147 CWE categories. Each task pairs repository snapshots immediately before and after a security fix. On the vulnerable snapshot, an agent receives only the CWE description and read-only terminal access and must return the affected fi...


自动采集于 2026-09-16

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录