[论文] Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities
论文概要
研究领域: cs.CR, cs.AI 作者: Md Erfan, Ahmed Ryan, Md Kamal Hossain Chowdhury 发布时间: 2026-07-21 arXiv: 2507.15483中文摘要
网联自动驾驶汽车(CAV)依赖相互连接的软件和硬件组件,包括传感器、电子控制单元、车载信息娱乐系统和远程信息处理单元,其中漏洞可能危及资产、用户和车辆运行。这些漏洞通常以纯文本形式记录在通用漏洞披露(CVE)数据库中;然而,安全从业者需要关于受影响资产、弱点类型和攻击行为的结构化信息,以有效缓解这些漏洞带来的风险。为此,我们评估了开源权重的大语言模型(LLM)为CAV相关CVE生成结构化威胁信息表达(STIX)的能力——STIX是一种著名的威胁信息结构化表示格式。我们构建了一个名为CAV-STIXGen的数据集,将CAV漏洞描述映射到STIX域对象(SDO)、STIX关系对象(SRO)、通用弱点枚举(CWE)和MITRE ATT&CK技术映射。使用该数据集,我们在各种提示策略和温度下评估了11个开源权重LLM(4B到120B参数)。单模型配置在SDO上达到0.94的F1分数,在SRO上达到0.63,在CWE映射上达到0.99,而完整的MITRE ATT&CK映射仍然具有挑战性。在多智能体设置中,Gemma-4-31B和Codestral-22B分别在SDO和SRO上达到0.91和0.43的F1分数。最后,我们分析CWE和MITRE ATT&CK的共现以识别CAV领域中的反复威胁模式,展示AI辅助的漏洞到STIX转换如何自动化威胁情报并优先处理交通安全防御。原文摘要
Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, Electronic Control Units, in-vehicle infotainment systems, and telematics units, where vulnerabilities can compromise assets, users, and vehicle operations. These vulnerabilities are commonly documented as plain text in the Common Vulnerabilities and Exposures (CVE) database; however, security practitioners require structured information about affected assets, types of weaknesses, and attack behaviors to effectively mitigate the risks from these vulnerabilities. To this end, we evaluate open-weight Large Language Models (LLMs) for generating Structured Threat Information Expression (STIX), a well-known structured format for representing threat information, for CAV-related CVEs. We construct a dataset called CAV-STIXGen that maps CAV vulnerability descriptions to STIX domain objects (SDO), STIX relationship objects (SRO), Common Weakness Enumeration (CWE), and MITRE ATT&CK techniques mappings. Using this dataset, we evaluated 11 open-weight LLMs (4B to 120B parameters) across various prompting strategies and temperatures. Single-model configurations achieve F1 scores of 0.94 for SDO, 0.63 for SRO, and 0.99 for CWE mapping, while complete MITRE ATT&CK mapping remains challenging. In a multi-agent setup, Gemma-4-31B and Codestral-22B achieve F1 scores of 0.91 for SDOs and 0.43 for SROs, respectively. Lastly, we analyze CWE and MITRE ATT&CK co-occurrences to identify recurring threat patterns in the CAV domain, demonstrating how AI-assisted vulnerability-to-STIX translation can automate threat intelligence and prioritize defense in transportation security.--- *自动采集于 2026-07-21*
#论文 #arXiv #CR #小凯
💬 讨论回复 (0)
推荐
🌟 智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。
🎁 领取 2000万 Tokens