论文概要
研究领域: NLP
作者: Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao
发布时间: 2026-10-05
arXiv: 2610.00328
中文摘要
结构化工具调用往往仅少数字段违反模式或执行契约即失败。重生成完整对象会扩大动作面并使重复修复难以审计。我们提出 ContractRL:契约约束顺序修复协议,将验证器引导的 JSON 修复建模为有界决策过程。每步策略观察候选对象、类型化验证器反馈、JSON 指针、不可变修复历史和剩余预算;契约派生动作掩码在确定性验证器执行转移前过滤格式错误或禁止的 RFC-6902 操作。我们为补丁、重试、弃权决策指定契约约束组相对目标,规范目标和语义标签保留在线状态外至轨迹冻结。相同验证器信息下,ContractRL 达 0.9362 语义成功率和 34.4 生成 token,Patch-SFT 为 0.9076/44.9,完整重生成为 0.9148/137.2(每种子 192 案例,五种子)。策略优化将监督版从 0.9186 提至 0.9375。独立三种子配对评估相对 Patch-SFT 语义差 +0.0396(95% CI [+0.0137, +0.0662],p=0.0039)。反馈、动作掩码、预算和模式漂移分析将增益连至局部化修正,对抗和多轮评估刻画剩余故障模式。
原文摘要
Structured tool calls often fail after only a small number of fields violate a schema or an execution contract. Regenerating the complete object enlarges the action surface and makes repeated repair difficult to audit. We introduce ContractRL, a contract-constrained sequential repair protocol that models verifier-guided JSON repair as a bounded decision process. At each step the policy observes the candidate, typed verifier feedback, JSON Pointer, immutable repair history, and remaining budget; a contract-derived action mask filters malformed or prohibited RFC-6902 operations before a deterministic validator performs the transition. We specify a contract-constrained group-relative objective for patch, retry, and abstention decisions while keeping canonical targets and semantic labels outsi...
自动采集于 2026-10-05
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。