论文概要
研究领域: ML
作者: Wooyoung Jung
发布时间: 2026-10-05
arXiv: 2610.00224
中文摘要
楼宇自动化系统越来越多地用 Brick、ASHRAE 223P 等本体表示为语义知识图谱,为 AI 应用提供机器可读基底。有前景的应用是将自然语言问题翻译为 SPARQL(text-to-SPARQL),让操作员通过语言智能体查询图谱,但进展受限于大规模基准稀缺。本文提出 Build2SPARQL,由知识图谱接地的流水线生成:SPARQL 查询完全由图遍历代码产生并验证,大模型只生成自然语言问题,使查询正确性与模型行为无关。流水线挖掘六类查询模式——线性链、分支、UNION、聚合、OPTIONAL、属性过滤,以五种词汇语域表述每个查询。应用于 201 个楼宇知识图谱(180 Brick、21 ASHRAE 223P),产出 6,136 个可执行 SPARQL 查询和 30,680 个问题。300 题双人验证:98.8% 语义保真、98.8% 自然度、84.0% 操作合理性。三个开源权重模型的检索增强评估将精确匹配从零样本 0.2–20% 提升到三样本检索 56–65%。
原文摘要
Building automation systems are increasingly represented as semantic knowledge graphs (KGs) using ontologies such as Brick and ASHRAE 223P, creating a machine-readable substrate for artificial-intelligence applications. One promising application is translating natural-language questions into SPARQL (text-to-SPARQL), which would let building operators query these graphs through language agents, but progress is limited by the scarcity of large natural-language/SPARQL benchmarks. This paper presents Build2SPARQL, a large-scale benchmark for building KGs generated by a KG-grounded pipeline: SPARQL queries are produced and validated entirely by graph-traversal code, while large language models generate only the natural-language questions, keeping query correctness independent of model behavior....
自动采集于 2026-10-05
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。