[论文] [论文] Metrics Failure in LLM-Based Code Vulnerability Repair: An E...

论文概要 研究领域: 安全 作者: Om Nepal, Sushant Aryal, Oluseyi Olukola, Nick Rahimi 发布时间: 2026-09-22 arXiv: 2609.26749

论文概要

研究领域: 安全 作者: Om Nepal, Sushant Aryal, Oluseyi Olukola, Nick Rahimi 发布时间: 2026-09-22 arXiv: 2609.26749

中文摘要

LLM 越来越多地用于 C/C++ 漏洞自动修复,编译率是常用进展代理指标。我们论证它对单函数漏洞修复科学上不可靠,并以五个对照实验支撑(Big-Vul 203 个漏洞函数、三个开源代码 LLM、三种提示策略)。编译率(i)对实质改进代码的干预几乎无响应;(ii)约 64% 编译失败不可归因于模型,且各模型间几乎不变;(iii)相同补丁仅改一个编译器标准标志就变动 1.8–2.7 倍且零回归;(iv)模型排序与参考相似度指标相反;(v)作优化目标会奖励"非修复"——编译器反馈回路提高编译率的同时与人类修复相似度下降,人工检查发现删除式与占位符式非修复。替代指标整体函数 CodeBLEU 同样失败:未修改的漏洞输入副本得分超过每个模型。我们提出 diff_F1——只给编辑区域打分的变化感知筛选:对空操作恰给零分,对部分删除式博弈补丁给近零分,仍认可真正的部分编辑,可作深入执行分析前的廉价筛选。结论:漏洞修复评估须变化感知、执行接地。

原文摘要

Large language models (LLMs) are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, and we support this with five controlled experiments over 203 vulnerable functions from Big-Vul, three open-source code LLMs (350M to 6.7B parameters), and three prompting strategies. Compile rate (i) barely responds to an intervention that substantially improves the generated code; (ii) is dominated by evaluation-harness and dataset artifacts rather than model quality, with about 64% of compile failures not attributable to the model, a share that is nearly invariant across models; (iii) shifts by 1.8 to 2.7 times on identical patches under a single compiler-standard flag, with zero regressions; (iv) ranks the three models in the opposite order to reference-similarity metrics; and (v) rewards non-repairs when used as an optimization target, since a compiler-feedback loop raises compile rate while similarity to the human fix falls, with manual inspection finding deletion- and placeholder-style non-repairs among the newly compiling outputs. The natural fallback, whole-function CodeBLEU, also fails: an unchanged copy of the vulnerable input outscores every model. We also examine diff_F1, a change-aware screen that scores only the edited region. It gives exactly zero credit to a no-op and near-zero credit to some, though not all, of the deletion-based gaming patches we observed, while still crediting genuine partial edits, so it may serve as a cheap screen before deeper, execution-based analysis. It is not a repair-quality metric, and we report where it falls short. Our findings argue for change-aware, execution-grounded evaluation of LLM-based vulnerability repair.


*自动采集于 2026-09-24*

#论文 #arXiv #安全 #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens