论文概要
研究领域: ML
作者: Ali ArjomandBigdeli, Jiawei Zhou, Stanley Bak
发布时间: 2026-09-17
arXiv: 2609.20752
中文摘要
反例生成(falsification)为信息物理系统(CPS)的形式化规约搜索反例。当规约以信号时序逻辑(STL)书写时,反例生成可表述为鲁棒度优化问题,传统上由黑盒搜索算法求解。与此同时,大语言模型(LLM)在与迭代式提示结合后,近来展现出惊人的优化器能力。本文将这两条线索连接起来,提出 LLM-Falsifier:一种通过最小化 STL 鲁棒度来证伪规约的 LLM 方法。超越通用的提示式优化,我们的关键思路是让 LLM 接触对语言模型而言很自然、却为标准数值优化器所缺失的语义信息:自然语言描述的输入输出名、输出轨迹,以及最小鲁棒度值的关键时刻见证(witness)。这些补充使鲁棒度搜索更智能、更省样本。在 ARCH-COMP 反例生成基准上,以找到反例所需平均仿真次数衡量,LLM-Falsifier 在 21 条规约中的 14 条上优于基于各类优化范式的现有工具——包括代理模型优化、贝叶斯优化与基于搜索的测试。
原文摘要
Falsification searches for counterexamples to formal specifications in cyber-physical systems (CPS). With specifications written in Signal Temporal Logic (STL), falsification can be formulated as a robustness optimization problem, traditionally tackled with black-box search algorithms. In parallel, large language models (LLMs) have recently emerged as surprisingly effective optimizers when coupled with iterative prompting. In this work, we connect these ideas and introduce LLM-Falsifier, an LLM-based approach that falsifies specifications by minimizing the STL robustness degree. Beyond generic prompt-based optimization, our key idea is to expose the LLM to semantic information that is natural for language models but absent from standard numerical optimizers, including natural-language inpu...
自动采集于 2026-09-20
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。