[论文] Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Build...

研究领域: ML 作者: Wooyoung Jung 发布时间: 2026-10-05 arXiv: 2610.00224

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Wooyoung Jung 发布时间: 2026-10-05 arXiv: 2610.00224

中文摘要

楼宇自动化系统越来越多地用 Brick、ASHRAE 223P 等本体表示为语义知识图谱,为 AI 应用提供机器可读基底。有前景的应用是将自然语言问题翻译为 SPARQL(text-to-SPARQL),让操作员通过语言智能体查询图谱,但进展受限于大规模基准稀缺。本文提出 Build2SPARQL,由知识图谱接地的流水线生成:SPARQL 查询完全由图遍历代码产生并验证,大模型只生成自然语言问题,使查询正确性与模型行为无关。流水线挖掘六类查询模式——线性链、分支、UNION、聚合、OPTIONAL、属性过滤,以五种词汇语域表述每个查询。应用于 201 个楼宇知识图谱(180 Brick、21 ASHRAE 223P),产出 6,136 个可执行 SPARQL 查询和 30,680 个问题。300 题双人验证:98.8% 语义保真、98.8% 自然度、84.0% 操作合理性。三个开源权重模型的检索增强评估将精确匹配从零样本 0.2–20% 提升到三样本检索 56–65%。

原文摘要

Building automation systems are increasingly represented as semantic knowledge graphs (KGs) using ontologies such as Brick and ASHRAE 223P, creating a machine-readable substrate for artificial-intelligence applications. One promising application is translating natural-language questions into SPARQL (text-to-SPARQL), which would let building operators query these graphs through language agents, but progress is limited by the scarcity of large natural-language/SPARQL benchmarks. This paper presents Build2SPARQL, a large-scale benchmark for building KGs generated by a KG-grounded pipeline: SPARQL queries are produced and validated entirely by graph-traversal code, while large language models generate only the natural-language questions, keeping query correctness independent of model behavior....


*自动采集于 2026-10-05*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens