Key points
This forum post reviews NLAH (Natural-Language Agent Harnesses) — a paper by Linyue Pan, Lexiao Zou, Shuo Guo, Jingchen Ni, and Hai-Tao Zheng (Tsinghua SIGS + Harbin Institute of Technology, Shenzhen), arXiv March 2026.
Paper: https://arxiv.org/abs/2603.25723 · Code: https://github.com/zjsxply/
The problem
Agent performance depends heavily on the harness — the execution framework around the model (multi-step reasoning, tool calls, state, recovery, verification, delegation). But harness logic is scattered across controller code, framework defaults, and undocumented runtime conventions. As LangChain put it: *"If you're not the Model, you're the Harness."* Two systems claiming to use the same model actually differ in tool interfaces, verification gates, and state carriers — making ablation impossible.
Core idea
- NLAH: an editable Markdown document describing task-run policy: phases (Plan → Execute → Verify → Recover → Finalize), roles (Solver / Verifier / Orchestrator), state rules, verification rules, failure recovery, stop conditions.
- IHR (Intelligent Harness Runtime): a shared runtime translating the NLAH document into actual agent calls, handoffs, state updates, verification gates, and artifact contracts.
- NLAH+IHR matches or exceeds code harnesses; on Live-SWE it beats the native code agent (73.0 vs 67.0) with shorter wall-clock time.
- Strategy compression: 60.1k → 2.9k tokens (Live-SWE, 20x); 10.5k → 0.8k (MHTBA, 13x). The reusable policy layer was only ~5% of the code but was previously inseparable.
- Mechanism audits: Artifact Contract = 1.000, Tool Call Success = 0.933, Failed Tool Continuation = 0.992 on Live-SWE — the runtime genuinely enforces the policy. Weakness: information handoff recall drops to 0.322 (Live-SWE) due to parent-child architecture.
- File-backed state: consistent gains (75.6 on Live-SWE, 58.3 on OSWorld).
- Self-evolution: best SWE score (78.8) but token-expensive.
- Multi-candidate search: agent calls jump 1.1 → 5.7, yet performance *drops* (73.0 → 71.4) — more search ≠ better harnessing.
- Context compression: no gain at all — compression loses information agents can't verify.
- Higher token consumption and call counts than code implementations (prototype runtime overhead).
- Low handoff recall in parent-child setups.
- Natural language can't faithfully represent mechanisms relying on hidden server-side state.
- Runtime charter may absorb behavior that should be attributed to the NLAH text (attribution pollution).
- Researchers: publish harness policies as readable .md documents; report module ablations and mechanism audits, not just aggregate scores.
- Engineers: separate policy (belongs in NLAH) from mechanism (belongs in code); extend AGENTS.md / CLAUDE.md / SKILL.md experience to the run-level harness.
- Framework authors: make LangChain/CrewAI/AutoGen defaults configurable as NLAH documents backed by an IHR-style shared runtime.
Natural language carries policy; code carries precise mechanisms (tool execution, parsing, sandboxing, logging). Four layers: Base Agent (LLM + terminal) ← fixed Runtime Policy ← swappable NLAH document ← Scripts & Adapters.
Results
| Benchmark | Code Harness | Prompted NLAH | NLAH + IHR | |------|------|--------|----------| | Live-SWE | 67.0 | 77.0 | 73.0 | | MHTBA (TB2) | 36.0 | 57.3 | 53.9 | | OSWorld | 47.1 | — | 46.3 |
Module ablations (the key contribution)
Why it matters
Harness design becomes a scientific, ablatable research object for the first time. The paper positions NLAH as a midpoint on a spectrum from restrictive code harnesses toward self-harnessing, where a model controls itself without an external layer.