[论文] QuoteBench: How Matched Scores Can Hide Command-Path Failures

论文概要 研究领域: ML 作者: Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang 发布时间: 2026-08-13 arXiv: 2608.13547

论文概要

研究领域: ML 作者: Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang 发布时间: 2026-08-13 arXiv: 2608.13547

中文摘要

LLM编码智能体通过可能序列化、包装和重新解析模型输出的接口发出Bash命令。仅靠匹配执行分数无法区分命令生成错误与生成后引入的失败。QuoteBench通过在14个事件衍生家族的56个一次性任务上进行精确最终状态验证来测量这一边界,跨越生成契约与执行传输,围绕一个故意未转义的添加解析器。在插值点进行转义可以复现每个重放回复的原始路径结果,因此在披露边界下的任何恢复必须来自模型改变其生成。在八种相同窗口配置中,通过添加解析器重放相同回复将成功率降低55.4至73.2个百分点;披露为六种配置恢复了30.4至60.7个百分点,对另外两种则恢复为零或略负。原始生成在前沿几乎已饱和;边界适应仍然是区分模型的因素。GPT-5.6-sol的匹配差距为-3.6分,隐藏了-64.3分的损害和+60.7分的补偿。部署配置重新排序模型:26个可比较对中有1个反转是明确的,另有4个处于单任务边缘。命令发出智能体的评估应报告模型配置、生成契约、执行路径、操作点和最终状态验证器,而不是将匹配分数视为模型的内在属性。

原文摘要

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation. Across eight same-window configurations, replaying the same reply through the added parser lowers success by 55.4 to 73.2 percentage points; disclosure recovers 30.4 to 60.7 points for six configurations, and zero or slightly negative for the other two. Raw generation is nearly saturated at the frontier; boundary adaptation is what still separates models. GPT-5.6-sol's matched gap of -3.6 points hides -64.3 points of damage and +60.7 points of compensation. The deployment configuration reorders models: one reversal among 26 comparable pairs is unambiguous and four more sit on single-task margins. Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.


*自动采集于 2026-08-15*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(1)

Q

QuoteBench(arXiv:2608.13547)抓的是 AI 代码生成里一个被严重低估的失败模式——「能写出对的逻辑,但被工具链解析坑死」。补几条:

① 56 个任务 / 14 个事故族这个设计值得称道:它不是泛泛测「代码对不对」,而是专门筛出那些「代码本身能跑、但交付物格式不符合下游解析器预期」的事故。未转义字符串、错误引号、缺逗号——这些在 LLM 输出里是高频错,在 human 代码里几乎不会发生。

② 未转义解析器导致成功率下降 55.4-73.2 个百分点,这个幅度比绝大多数「模型能力」差异都大。换句话说:很多模型在代码 benchmark 上的分差,可能还没「它有没有把 JSON 转义对」这件事影响大。这是把「模型评测」和「工具链鲁棒性评测」混为一谈的代价。

③ GPT-5.6-sol 隐藏 -64.3 / +60.7 的分化很有意思:同一个模型,在「隐藏」条件下(不给它看完整上下文?)暴跌,在另一种条件下暴涨。这说明模型的「事故倾向」高度依赖 prompt 结构和可见信息,不是固定属性。对工程来说意味着:同样的模型,换种包装方式事故率能差一个数量级。

④ 边界适应(boundary adaptation)拉开模型差距这个点,本质是说「能不能感知自己在和脆弱解析器打交道」是一种独立能力。会自我校正的模型在 QuoteBench 上明显更稳,这和「元认知」是 2026 年模型分化的暗线完全吻合。

⑤ 下一根钉子在「benchmark 本身会不会被快速过拟合」。这类高度结构化的事故族,一旦公开,模型训练时会悄悄把对应格式偏好烧进去,两年后重测可能全部满分——评测的有效性寿命很短。

收尾:QuoteBench 真正提醒业界的是——在 Agent 化代码交付里,「语法正确」和「可被消费」之间隔着一整条工具链鸿沟,而这条鸿沟从来不在任何主流 code benchmark 的视野里。下一根最该盯的钉子是它能不能推动下游框架(比如把 LLM 输出先过一层 schema 校验再进解析器)变成默认实践,而不是只停留在论文里的红色下降条。

暂无表态

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens