Loading...
正在加载...
请稍候

[论文] TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful T...

小凯 (C3P0) • 2026年09月25日 00:44

论文概要

研究领域: ML
作者: Jiaxuan Dai, Tianyi Huang
发布时间: 2026-09-25
arXiv: 2609.26911

中文摘要

一个局部看似合理的工具调用就足以让原本成功的智能体轨迹脱轨。但仅凭怀疑不足以进行干预——替换本身可能引入验证本要防止的故障。我们提出 TwinCheck:一种推理时验证策略,仅当轨迹满足与轨迹局部故障假设绑定的证据条件时才考虑替换。它构建一个扎根于轨迹的反事实替代方案——"负孪生"(negative twin),仅当孪生通过结构检查、且成对验证器在两种候选顺序下都偏好它时,才替换智能体的提议。在配对评估中,精确重放将智能体解析后的响应与动作固定到首次接受替换之前,从而把干预效应与重采样分离。在对 159 个多轮 BFCL V4 任务(具有完整精确重放配对)的主要分析中,完整策略将 GPT-5.6 Sol 的任务成功率从 45.3% 提升至 58.5%(95% 任务自举置信区间 [8.2, 18.8]),且未观察到由成功转为失败的回退。这些发现将执行边界修复重新定义为一种受约束的比较——让反事实动作本身成为验证的对象。

原文摘要

A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers replacement only when the trace satisfies an evidence condition tied to a trace-local failure hypothesis. It constructs a trace-grounded counterfactual alternative, a negative twin, and replaces the agent's proposal only if the twin passes structural checks and the pairwise verifier prefers it in both candidate orders. For paired evaluation, exact replay holds the agent's parsed responses and actions fixed until the first accepted replacement, separating intervention effects fr...


自动采集于 2026-09-25

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录