English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

QuoteBench: How Matched Scores Can Hide Command-Path Failures in LLM Coding Agents

Forum topic · 小凯 · 2026-08-15

Summary

QuoteBench (arXiv:2608.13547) studies a blind spot in evaluating LLM coding agents that issue Bash commands through interfaces that serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. The benchmark uses exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with an execution transport around a deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation. Across eight same-window configurations, replaying replies through the added parser lowers success by 55.4 to 73.2 percentage points; disclosure recovers 30.4 to 60.7 points for six configurations and zero or slightly negative for the other two. Raw generation is nearly saturated at the frontier; boundary adaptation separates models. GPT-5.6-sol's matched gap of -3.6 points hides -64.3 points of damage and +60.7 points of compensation, and deployment configuration reorders model rankings. The authors argue evaluations should report model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.

QuoteBench: How Matched Scores Can Hide Command-Path Failures

Field: ML Authors: Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang Published: 2026-08-13 arXiv: 2608.13547

Abstract

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser.

Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation.

Key findings

  • Across eight same-window configurations, replaying the same reply through the added parser lowers success by 55.4 to 73.2 percentage points.
  • Disclosure recovers 30.4 to 60.7 points for six configurations, and zero or slightly negative for the other two.
  • Raw generation is nearly saturated at the frontier; boundary adaptation is what still separates models.
  • GPT-5.6-sol's matched gap of -3.6 points hides -64.3 points of damage and +60.7 points of compensation.
  • The deployment configuration reorders models: one reversal among 26 comparable pairs is unambiguous, and four more sit on single-task margins.

Recommendation

Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.

---

Links: arXiv:2608.13547

Tags

#llm#coding-agents#benchmark#evaluation#bash#quotebench#arxiv#machine-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633500