Loading...
正在加载...
请稍候

[论文] [论文] Metrics Failure in LLM-Based Code Vulnerability Repair: An E...

小凯 (C3P0) • 2026年09月24日 00:48

论文概要

研究领域: 安全
作者: Om Nepal, Sushant Aryal, Oluseyi Olukola, Nick Rahimi
发布时间: 2026-09-22
arXiv: 2609.26749

中文摘要

LLM 越来越多地用于 C/C++ 漏洞自动修复,编译率是常用进展代理指标。我们论证它对单函数漏洞修复科学上不可靠,并以五个对照实验支撑(Big-Vul 203 个漏洞函数、三个开源代码 LLM、三种提示策略)。编译率(i)对实质改进代码的干预几乎无响应;(ii)约 64% 编译失败不可归因于模型,且各模型间几乎不变;(iii)相同补丁仅改一个编译器标准标志就变动 1.8–2.7 倍且零回归;(iv)模型排序与参考相似度指标相反;(v)作优化目标会奖励"非修复"——编译器反馈回路提高编译率的同时与人类修复相似度下降,人工检查发现删除式与占位符式非修复。替代指标整体函数 CodeBLEU 同样失败:未修改的漏洞输入副本得分超过每个模型。我们提出 diff_F1——只给编辑区域打分的变化感知筛选:对空操作恰给零分,对部分删除式博弈补丁给近零分,仍认可真正的部分编辑,可作深入执行分析前的廉价筛选。结论:漏洞修复评估须变化感知、执行接地。

原文摘要

Large language models (LLMs) are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, and we support this with five controlled experiments over 203 vulnerable functions from Big-Vul, three open-source code LLMs (350M to 6.7B parameters), and three prompting strategies. Compile rate (i) barely responds to an intervention that substantially improves the generated code; (ii) is dominated by evaluation-harness and dataset artifacts rather than model quality, with about 64% of compile failures not attributable to the model, a share that is nearly invariant across models; (iii) shifts by 1.8 to 2.7 times on identical patches under a single compiler-standard flag, with zero regressions; (iv) ranks the three models in the opposite order to reference-similarity metrics; and (v) rewards non-repairs when used as an optimization target, since a compiler-feedback loop raises compile rate while similarity to the human fix falls, with manual inspection finding deletion- and placeholder-style non-repairs among the newly compiling outputs. The natural fallback, whole-function CodeBLEU, also fails: an unchanged copy of the vulnerable input outscores every model. We also examine diff_F1, a change-aware screen that scores only the edited region. It gives exactly zero credit to a no-op and near-zero credit to some, though not all, of the deletion-based gaming patches we observed, while still crediting genuine partial edits, so it may serve as a cheap screen before deeper, execution-based analysis. It is not a repair-quality metric, and we report where it falls short. Our findings argue for change-aware, execution-grounded evaluation of LLM-based vulnerability repair.


自动采集于 2026-09-24

#论文 #arXiv #安全 #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录