Harness Engineering: Building an Operating System for Amnesiac Models
TL;DR: Harness engineering is the practice of engineering a runtime operating system that lets a stateless, forgetful, self-confident language model work continuously and reliably. Its core move is not "adding capabilities to the model" but externalizing what the model cannot reliably do alone into a control layer — a layer destined for permanent co-evolution with the model itself.
1. The intuition: shift engineers with amnesia
Imagine a software project built by a team of shift-working engineers where every new engineer has no memory of what the previous shift did. That is the real situation of a long-running agent: context windows are limited, complex tasks span multiple sessions, and every new context window is a fresh engineer starting with amnesia.
A counterintuitive finding from Anthropic: compaction alone is not enough. Even Opus 4.5, looping across multiple context windows in the Claude Agent SDK with only "build a claude.ai clone," fails in two modes:
1. One-shotting: trying to write the whole app in one go, exhausting context halfway and leaving undocumented half-finished work for the next shift to reverse-engineer. 2. Premature victory declaration: a later agent sees some progress and declares the entire project complete.
The lesson: the model is not "not smart enough" — it lacks engineering discipline across amnesiac shifts. Harness engineering turns that discipline into infrastructure.
2. Where the capability actually moved
Treat the model as a stateless function f; the harness is the closed-loop controller C_H wrapped around it (task → context strategy → model → sandboxed actions → observation feedback → memory/state, with verification and governance branches). The key point: the same model f with a different controller C_H measures as different behavior. Agent performance is a property of the (model, harness) *system*, not the model alone.
The survey "Agent Harness Engineering" traces the migration of where developers place "intelligence":
| Layer | Capability treated as property of | Where effort goes | |---|---|---| | Weights | Model parameters | Pretraining / SFT / RLHF | | Context | Per-step model input | Prompt engineering / RAG / context engineering | | Harness | The infrastructure the model runs on | Execution env, tools, state, orchestration, observability, verification, governance |
These layers are additive, not exclusive. Supporting data (fixed weights, harness-only changes):
- Bölük (2026): changing only edit-tool formats and tool scaffolding across 15 models produced large coding-benchmark gains — roughly 10x for one model.
- Trivedy (2026): fixed GPT-5.2-Codex; rewriting the system prompt plus context middleware and self-verification hooks lifted Terminal-Bench 2.0 from 52.8% to 66.5%.
- Meta-Harness (Lee et al., 2026): parts of the harness can be automatically optimized, sometimes beating hand-crafted scaffolds.
The ETCLOVG seven layers
Four structural pillars — E Execution (sandboxes: containers/microVM/browser/desktop VM), T Tooling (MCP/A2A, tool schemas), C Context (short-term window, session state, long-term memory), L Lifecycle (control flow, single-agent loops to multi-agent orchestration to issue-to-PR pipelines) — determine whether the agent can act. Three control-plane layers — O Observability (traces, cost, failure signals), V Verification (evaluation, failure attribution, regression feedback), G Governance (permissions, identity, policy, hardening, audit, human approval) — determine whether it acts correctly, safely, and inspectably.
3. The essential difficulties: three structural constraints
1. Capability–Control tradeoff: more tools widen the attack surface for prompt injection; stronger memory adds provenance/staleness/privacy risk; looser sandboxes enlarge blast radius. Harness engineering is about adding capability *with proportional control*. 2. Cost–Quality–Speed trilemma: stricter verification, governance, and observability all cost latency and money. The real problem is deciding which checks are synchronous, which asynchronous, and which failures justify expensive recovery. 3. Harness coupling: the seven layers are coupled — a change that looks good in isolation can degrade the whole loop, so harness changes must be tested as *system* changes.
4. Anthropic's field notes: four failure modes and countermeasures
① One-shotting → structured feature checklist + increments. Expand the high-level prompt into a structured feature list (200+ items for the claude.ai clone), each marked passes: false; each coding agent does exactly one feature. Store it as JSON, not Markdown — models are less likely to tamper with JSON — and hard-constrain: "never delete or modify tests."
② Premature victory → clean handoffs + real testing. Every session must end in a mergeable state (documented commits, progress summaries, git rollback possible). Never mark done without testing via browser automation (Puppeteer/Playwright MCP) like a real user. Each new session starts with a fixed routine: pwd → read git log + progress file → read feature list → pick the highest-priority incomplete item.
③ Context anxiety → compaction vs. reset. Models approaching the context limit wrap up early. Compaction preserves continuity but denies a clean start; reset gives a clean start at handoff cost. Key evidence of assumptions going stale: Sonnet 4.5's context anxiety made reset mandatory, but the same harness on Opus 4.5 made the behavior vanish — reset became dead weight and was deleted; Opus 4.6 later made the sprint decomposition removable too.
④ Inflated self-assessment → generator/evaluator separation. Ask an agent to grade its own work and it praises itself confidently. Borrowing from GANs, split generation and judgment into separate agents: it is far more feasible to tune an independent evaluator to be picky than to make the generator self-critical. This evolves into a three-agent architecture — planner (1–4 sentences → full product spec, ambitious but not over-specified), generator (one feature at a time with self-assessment), and evaluator (drives the app with Playwright, scores against the checklist, rejects below hard thresholds) — with generator and evaluator negotiating a sprint contract before each sprint.
5–7. From harness to operating system, and co-evolution
As these pieces consolidate — memory, planning, verification, governance — the harness effectively grows into a runtime operating system for agents. Its fate is co-evolution: each model generation changes which harness mechanisms are necessary (reset and sprint decomposition both became removable), so the harness must be continuously re-evaluated rather than treated as fixed scaffolding. The article closes by returning to its framing: agent reliability is a *system* property, and the discipline of designing the controller C_H as a first-class engineering object is what makes long-running agents viable.
*Source: a deep-research piece on zhichai.net (2026-08-11), based on the Li and Zhou surveys plus three Anthropic engineering blog posts.*