论文概要
研究领域: NLP
作者: Dain Kim, Eungi Cho, Kyumin Kim, Shinyeong Noh, Kyuseong Lim
发布时间: 2026-09-04
arXiv: 2609.05395
中文摘要
随着数据主权法规日益要求公共机构部署可串联多个实时政府API调用的开源本地LLM智能体,开源模型在多步工具调用场景下始终表现不佳,且缺乏衡量这一差距的基准测试。本文提出韩国开放公共API基准(KOPA-Bench),包含145个真实任务。为弥合差距,我们提出EDGE——一种基于实时执行的动态图数据合成方法,构建工具输出如何馈入另一工具输入的图结构,仅保留对真实API调用成功的链接,并遍历这些验证链接合成可执行的多步轨迹。通过在生成数据集上GRPO微调,我们的9B模型在KOPA-Bench和BFCL基准上均大幅接近同系列未微调的27B模型。
原文摘要
Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this multi-step setting, and no existing benchmark measures the gap. We introduce the Korean Open Public API Benchmark (KOPA-Bench), comprising 145 real-world tasks. To close this gap, we present EDGE, an Execution-grounded Dynamic Graph for tool-calling data synthEsis driven by live execution. EDGE builds a graph of how each tool's output can feed another's input, keeps only the links that succeed when actually called against the live APIs, and traverses these verified links to synthesize executable multi-step trajectories. Fine-tuned via GRPO on the resulting ...
自动采集于 2026-09-09
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。