> 📌 This is a GEO-optimized version of the original topic.
> One-line takeaway: An agent fixed a bug, tests passed; after another revision round, tests still passed—but the original bug was back. A 2026 arXiv paper, *Looping Is Not Reliability*, quantifies this precisely.
Key points
- Correctness is not an absorbing state. Across 900 forced-revision trajectories (30 HumanEval problems × 5 seeds × 3 rounds, 7B agent), *ever-correct* rose from 82.0% to 85.3%, but *current correctness* dropped from 82.0% to 69.3% by round 3. In total, 16.0% of trajectories produced a correct patch and then lost it. With current-only evidence, attrition fell to 8.0% but remained non-zero.
- Stale evidence resurrects fixed bugs. In the 14B replication starting from already-correct states, stale traces re-broke fixed bugs in 34/135 cases (25.2%) vs. only 4/135 (3.0%) with current traces—a 22.2 pp gap, 95% CI [8.9, 37.0], Holm-corrected p=0.0337.
- Evidence helps repair but hurts retention. With current traces, 113/135 incorrect starts were fixed (83.7%); stale traces fixed 105/135 (77.8%); no evidence fixed 97/135 (71.9%).
- Verifier quality ≠ verifier independence. Cross-family validation (Qwen 7B/14B, DeepSeek 6.7B) shows a strong verifier can still share blind spots with the coder, so cross-family diversity does not guarantee independent judgment inside a loop.
- Evaluation blind spot: benchmarks like SWE-bench and HumanEval measure only final-version correctness, missing patches that were once correct and later regressed (85.3% ever-correct vs 69.3% current correctness).
- Architecture implications: generate-test-revise loops provide no reliability guarantee by themselves; without retention mechanisms and stop rules, iteration can diverge rather than converge.
- Granularity principle: optimization must target whole trajectories, not single calls—"tests passed once" ≠ "bug fixed on the trajectory level."
- Only 30 entry-level HumanEval problems; real SWE-bench tasks are far harder.
- Forced revision is artificially triggered; an adaptive 540-rollout policy experiment removed correct-start harm but reduced wrong-start repair and failed a joint criterion.
- Repository-level experiments (24 real bugs, 4 coder stacks) hit a floor effect.
- StateSeal is a reference implementation without production-scale data.
Proposed solution: StateSeal
The paper offers StateSeal, an evidence-bound typed loop contract (a reference implementation, not a new agent framework):
1. State-bound evidence — every test result is bound to a code-state hash; stale evidence auto-expires. 2. Typed revision actions — revisions have types with distinct admission/retention rules. 3. Last-known-good checkpoint — correct versions are saved and can be rolled back to. 4. Risk- and dependence-aware stopping — stop rules weigh the risk of further revisions. 5. Bundled admission — acceptance considers retention effects, not just the current test run.
Why it matters
Honest limitations
FAQ
Who is this for? Practitioners, researchers, and students interested in AI, machine learning, and coding agents.
Core takeaway? Iterative revision does not preserve correctness; stale evidence significantly re-breaks fixed bugs, so reliability needs state-bound evidence, checkpoints, and stop rules.
Open source? The paper mentions a StateSeal reference implementation; no public link on the arXiv page yet.
> Paper: arXiv:2607.24604