论文概要
研究领域: NLP
作者: Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer
发布时间: 2026-10-01
arXiv: 2610.02206
中文摘要
LLM 正越来越多地应用于网络安全工作流,需要将分析师意图转化为工具调用。然而现有评估聚焦于知识型测试或端到端智能体任务,并不直接衡量 LLM 为真实安全工具生成可执行命令的能力。这一缺口至关重要——网络安全操作依赖严格的命令行界面,微小的语法错误、错误的参数-值绑定或参数顺序错误都可能导致执行失败。我们提出 KaliBench,一个面向 Kali Linux 自然语言到 CLI 翻译的细粒度基准和数据集,包含 8504 条查询-命令对,覆盖 1642 个工具、23 个能力维度和 5 个安全阶段。KaliBench 通过基于文档的流水线构建,具备确定性规范化与别名感知评估,可精确且可复现地评估工具选择和参数构建。为确保语义正确性和实际可执行性,我们开发了结合 LLM 验证、沙箱终端执行和人工优化的多阶段验证流水线。基于这些细粒度确定性信号,KaliBench 还支持免运行时的可验证奖励训练。在三种评估模式和 24 种配置下,无开源模型在不受限设置中超过 42% 的精确命令准确率,凸显了在缺乏工具提示时准确使用 CLI 安全工具的难度。我们进一步表明,基于 KaliBench 的监督微调和可验证奖励强化学习可显著提升 8B 模型,使其达到与 685B MoE 模型相当的性能。
原文摘要
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is cons...
自动采集于 2026-10-03
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。