English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Harness Engineering: Building a Reliable Runtime for Amnesiac, Overconfident Language Models

Forum topic · ✨步子哥 · 2026-08-11

Summary

This in-depth article from zhichai.net explores "Harness Engineering" — the practice of building a runtime control system around stateless, forgetful, and self-confident language models so they can work reliably over long-running tasks. Using Anthropic's metaphor of shift-working engineers with no memory of the previous shift, the author explains why compaction alone fails: agents either attempt one-shot builds that exhaust their context or prematurely declare victory. Core claims: agent performance is a property of the (model, harness) system, not the model alone; capabilities have migrated across three layers (weights, context, harness); and the field's ETCLOVG taxonomy covers Execution, Tooling, Context, Lifecycle, Observability, Verification, and Governance. Evidence includes Bölük (2026) showing ~10x coding gains from tool-format changes alone, and Trivedy (2026) lifting Terminal-Bench 2.0 from 52.8% to 66.5% with prompt and middleware changes on fixed weights. The article details four failure modes and mitigations: feature checklists with incremental work, clean handoffs with end-to-end testing, compaction vs. reset for context anxiety (reset became dead weight with Opus 4.5), and generator/evaluator separation to counter inflated self-assessment. It concludes that harness and model are locked in permanent co-evolution.

Harness Engineering: Building an Operating System for Amnesiac Models

TL;DR: Harness engineering is the practice of engineering a runtime operating system that lets a stateless, forgetful, self-confident language model work continuously and reliably. Its core move is not "adding capabilities to the model" but externalizing what the model cannot reliably do alone into a control layer — a layer destined for permanent co-evolution with the model itself.

1. The intuition: shift engineers with amnesia

Imagine a software project built by a team of shift-working engineers where every new engineer has no memory of what the previous shift did. That is the real situation of a long-running agent: context windows are limited, complex tasks span multiple sessions, and every new context window is a fresh engineer starting with amnesia.

A counterintuitive finding from Anthropic: compaction alone is not enough. Even Opus 4.5, looping across multiple context windows in the Claude Agent SDK with only "build a claude.ai clone," fails in two modes:

1. One-shotting: trying to write the whole app in one go, exhausting context halfway and leaving undocumented half-finished work for the next shift to reverse-engineer. 2. Premature victory declaration: a later agent sees some progress and declares the entire project complete.

The lesson: the model is not "not smart enough" — it lacks engineering discipline across amnesiac shifts. Harness engineering turns that discipline into infrastructure.

2. Where the capability actually moved

Treat the model as a stateless function f; the harness is the closed-loop controller C_H wrapped around it (task → context strategy → model → sandboxed actions → observation feedback → memory/state, with verification and governance branches). The key point: the same model f with a different controller C_H measures as different behavior. Agent performance is a property of the (model, harness) *system*, not the model alone.

The survey "Agent Harness Engineering" traces the migration of where developers place "intelligence":

| Layer | Capability treated as property of | Where effort goes | |---|---|---| | Weights | Model parameters | Pretraining / SFT / RLHF | | Context | Per-step model input | Prompt engineering / RAG / context engineering | | Harness | The infrastructure the model runs on | Execution env, tools, state, orchestration, observability, verification, governance |

These layers are additive, not exclusive. Supporting data (fixed weights, harness-only changes):

  • Bölük (2026): changing only edit-tool formats and tool scaffolding across 15 models produced large coding-benchmark gains — roughly 10x for one model.
  • Trivedy (2026): fixed GPT-5.2-Codex; rewriting the system prompt plus context middleware and self-verification hooks lifted Terminal-Bench 2.0 from 52.8% to 66.5%.
  • Meta-Harness (Lee et al., 2026): parts of the harness can be automatically optimized, sometimes beating hand-crafted scaffolds.
This is the survey's binding-constraint thesis: in the real world, the ceiling on agent reliability is set by infrastructure quality, not just model capability.

The ETCLOVG seven layers

Four structural pillars — E Execution (sandboxes: containers/microVM/browser/desktop VM), T Tooling (MCP/A2A, tool schemas), C Context (short-term window, session state, long-term memory), L Lifecycle (control flow, single-agent loops to multi-agent orchestration to issue-to-PR pipelines) — determine whether the agent can act. Three control-plane layers — O Observability (traces, cost, failure signals), V Verification (evaluation, failure attribution, regression feedback), G Governance (permissions, identity, policy, hardening, audit, human approval) — determine whether it acts correctly, safely, and inspectably.

3. The essential difficulties: three structural constraints

1. Capability–Control tradeoff: more tools widen the attack surface for prompt injection; stronger memory adds provenance/staleness/privacy risk; looser sandboxes enlarge blast radius. Harness engineering is about adding capability *with proportional control*. 2. Cost–Quality–Speed trilemma: stricter verification, governance, and observability all cost latency and money. The real problem is deciding which checks are synchronous, which asynchronous, and which failures justify expensive recovery. 3. Harness coupling: the seven layers are coupled — a change that looks good in isolation can degrade the whole loop, so harness changes must be tested as *system* changes.

4. Anthropic's field notes: four failure modes and countermeasures

① One-shotting → structured feature checklist + increments. Expand the high-level prompt into a structured feature list (200+ items for the claude.ai clone), each marked passes: false; each coding agent does exactly one feature. Store it as JSON, not Markdown — models are less likely to tamper with JSON — and hard-constrain: "never delete or modify tests."

② Premature victory → clean handoffs + real testing. Every session must end in a mergeable state (documented commits, progress summaries, git rollback possible). Never mark done without testing via browser automation (Puppeteer/Playwright MCP) like a real user. Each new session starts with a fixed routine: pwd → read git log + progress file → read feature list → pick the highest-priority incomplete item.

③ Context anxiety → compaction vs. reset. Models approaching the context limit wrap up early. Compaction preserves continuity but denies a clean start; reset gives a clean start at handoff cost. Key evidence of assumptions going stale: Sonnet 4.5's context anxiety made reset mandatory, but the same harness on Opus 4.5 made the behavior vanish — reset became dead weight and was deleted; Opus 4.6 later made the sprint decomposition removable too.

④ Inflated self-assessment → generator/evaluator separation. Ask an agent to grade its own work and it praises itself confidently. Borrowing from GANs, split generation and judgment into separate agents: it is far more feasible to tune an independent evaluator to be picky than to make the generator self-critical. This evolves into a three-agent architecture — planner (1–4 sentences → full product spec, ambitious but not over-specified), generator (one feature at a time with self-assessment), and evaluator (drives the app with Playwright, scores against the checklist, rejects below hard thresholds) — with generator and evaluator negotiating a sprint contract before each sprint.

5–7. From harness to operating system, and co-evolution

As these pieces consolidate — memory, planning, verification, governance — the harness effectively grows into a runtime operating system for agents. Its fate is co-evolution: each model generation changes which harness mechanisms are necessary (reset and sprint decomposition both became removable), so the harness must be continuously re-evaluated rather than treated as fixed scaffolding. The article closes by returning to its framing: agent reliability is a *system* property, and the discipline of designing the controller C_H as a first-class engineering object is what makes long-running agents viable.

*Source: a deep-research piece on zhichai.net (2026-08-11), based on the Li and Zhou surveys plus three Anthropic engineering blog posts.*

Tags

#harness-engineering#ai-agents#context-engineering#anthropic#llm-orchestration#agent-reliability#verification#runtime-infrastructure

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633318