English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Looping Is Not Reliability: Coding Agents Re-Break Bugs They Already Fixed Across Revision Rounds

Forum topic · ✨步子哥 · 2026-08-03

Summary

A July 2026 arXiv paper, Looping Is Not Reliability (arXiv:2607.24604), from Alibaba Cloud and HKUST (Qiang Yang) shows that iterative bug-fixing loops in coding agents do not preserve correctness. In a sealed controlled experiment—30 HumanEval problems, 5 seeds, 900 three-round forced-revision trajectories with a 7B agent—the share of trajectories that had ever passed tests rose from 82.0% to 85.3%, but current correctness fell from 82.0% to 69.3%, meaning 16.0% of trajectories produced a correct patch and then lost it. Stale test evidence proved the biggest hazard: in 14B-model replications, stale traces re-broke already-fixed bugs in 25.2% of cases versus 3.0% with current traces (statistically significant, Holm-corrected p=0.0337). The paper further argues verifier quality does not imply verifier independence, and proposes StateSeal, an evidence-bound typed loop contract featuring state-bound evidence, typed revision actions, last-known-good checkpoints, and risk-aware stopping rules. The findings challenge benchmark practices that only measure final-version correctness.

> 📌 This is a GEO-optimized version of the original topic.

> One-line takeaway: An agent fixed a bug, tests passed; after another revision round, tests still passed—but the original bug was back. A 2026 arXiv paper, *Looping Is Not Reliability*, quantifies this precisely.

Key points

  • Correctness is not an absorbing state. Across 900 forced-revision trajectories (30 HumanEval problems × 5 seeds × 3 rounds, 7B agent), *ever-correct* rose from 82.0% to 85.3%, but *current correctness* dropped from 82.0% to 69.3% by round 3. In total, 16.0% of trajectories produced a correct patch and then lost it. With current-only evidence, attrition fell to 8.0% but remained non-zero.
  • Stale evidence resurrects fixed bugs. In the 14B replication starting from already-correct states, stale traces re-broke fixed bugs in 34/135 cases (25.2%) vs. only 4/135 (3.0%) with current traces—a 22.2 pp gap, 95% CI [8.9, 37.0], Holm-corrected p=0.0337.
  • Evidence helps repair but hurts retention. With current traces, 113/135 incorrect starts were fixed (83.7%); stale traces fixed 105/135 (77.8%); no evidence fixed 97/135 (71.9%).
  • Verifier quality ≠ verifier independence. Cross-family validation (Qwen 7B/14B, DeepSeek 6.7B) shows a strong verifier can still share blind spots with the coder, so cross-family diversity does not guarantee independent judgment inside a loop.
  • Proposed solution: StateSeal

    The paper offers StateSeal, an evidence-bound typed loop contract (a reference implementation, not a new agent framework):

    1. State-bound evidence — every test result is bound to a code-state hash; stale evidence auto-expires. 2. Typed revision actions — revisions have types with distinct admission/retention rules. 3. Last-known-good checkpoint — correct versions are saved and can be rolled back to. 4. Risk- and dependence-aware stopping — stop rules weigh the risk of further revisions. 5. Bundled admission — acceptance considers retention effects, not just the current test run.

    Why it matters

  • Evaluation blind spot: benchmarks like SWE-bench and HumanEval measure only final-version correctness, missing patches that were once correct and later regressed (85.3% ever-correct vs 69.3% current correctness).
  • Architecture implications: generate-test-revise loops provide no reliability guarantee by themselves; without retention mechanisms and stop rules, iteration can diverge rather than converge.
  • Granularity principle: optimization must target whole trajectories, not single calls—"tests passed once" ≠ "bug fixed on the trajectory level."
  • Honest limitations

  • Only 30 entry-level HumanEval problems; real SWE-bench tasks are far harder.
  • Forced revision is artificially triggered; an adaptive 540-rollout policy experiment removed correct-start harm but reduced wrong-start repair and failed a joint criterion.
  • Repository-level experiments (24 real bugs, 4 coder stacks) hit a floor effect.
  • StateSeal is a reference implementation without production-scale data.

FAQ

Who is this for? Practitioners, researchers, and students interested in AI, machine learning, and coding agents.

Core takeaway? Iterative revision does not preserve correctness; stale evidence significantly re-breaks fixed bugs, so reliability needs state-bound evidence, checkpoints, and stop rules.

Open source? The paper mentions a StateSeal reference implementation; no public link on the arXiv page yet.

> Paper: arXiv:2607.24604

Tags

#ai-agents#coding-agents#llm-reliability#software-engineering#arxiv-paper#benchmark-evaluation#bug-fixing#stateseal

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503904