English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

NLAH: Turning Agent Harnesses from Code into Readable Markdown Documents

Forum topic · 小凯 · 2026-06-16

Summary

A Chinese tech forum post introduces NLAH (Natural-Language Agent Harnesses), a paper from Tsinghua University (Shenzhen) and Harbin Institute of Technology (Shenzhen) on arXiv:2603.25723. The work extracts agent harness logic—phases, roles, state rules, verification, recovery, stop conditions—out of scattered Python code and framework defaults into an editable Markdown policy document, executed by a shared runtime called IHR (Intelligent Harness Runtime). Evaluated on Live-SWE, Terminal-Bench 2.0 (MHTBA), and OSWorld, NLAH+IHR matches or exceeds code harnesses (73.0 vs 67.0 on Live-SWE; 53.9 vs 36.0 on MHTBA) while compressing strategy definitions by 13-20x. Mechanism audits confirm real contract enforcement and tool-mediated execution. Module ablations show file-backed state and self-evolution help, while multi-candidate search and context compression hurt or add nothing. The post also covers limitations (token overhead, handoff recall, attribution pollution) and implications for researchers, engineers, and framework authors.

Key points

This forum post reviews NLAH (Natural-Language Agent Harnesses) — a paper by Linyue Pan, Lexiao Zou, Shuo Guo, Jingchen Ni, and Hai-Tao Zheng (Tsinghua SIGS + Harbin Institute of Technology, Shenzhen), arXiv March 2026.

Paper: https://arxiv.org/abs/2603.25723 · Code: https://github.com/zjsxply/

The problem

Agent performance depends heavily on the harness — the execution framework around the model (multi-step reasoning, tool calls, state, recovery, verification, delegation). But harness logic is scattered across controller code, framework defaults, and undocumented runtime conventions. As LangChain put it: *"If you're not the Model, you're the Harness."* Two systems claiming to use the same model actually differ in tool interfaces, verification gates, and state carriers — making ablation impossible.

Core idea

  • NLAH: an editable Markdown document describing task-run policy: phases (Plan → Execute → Verify → Recover → Finalize), roles (Solver / Verifier / Orchestrator), state rules, verification rules, failure recovery, stop conditions.
  • IHR (Intelligent Harness Runtime): a shared runtime translating the NLAH document into actual agent calls, handoffs, state updates, verification gates, and artifact contracts.
  • Natural language carries policy; code carries precise mechanisms (tool execution, parsing, sandboxing, logging). Four layers: Base Agent (LLM + terminal) ← fixed Runtime Policy ← swappable NLAH document ← Scripts & Adapters.

    Results

    | Benchmark | Code Harness | Prompted NLAH | NLAH + IHR | |------|------|--------|----------| | Live-SWE | 67.0 | 77.0 | 73.0 | | MHTBA (TB2) | 36.0 | 57.3 | 53.9 | | OSWorld | 47.1 | — | 46.3 |

  • NLAH+IHR matches or exceeds code harnesses; on Live-SWE it beats the native code agent (73.0 vs 67.0) with shorter wall-clock time.
  • Strategy compression: 60.1k → 2.9k tokens (Live-SWE, 20x); 10.5k → 0.8k (MHTBA, 13x). The reusable policy layer was only ~5% of the code but was previously inseparable.
  • Mechanism audits: Artifact Contract = 1.000, Tool Call Success = 0.933, Failed Tool Continuation = 0.992 on Live-SWE — the runtime genuinely enforces the policy. Weakness: information handoff recall drops to 0.322 (Live-SWE) due to parent-child architecture.
  • Module ablations (the key contribution)

  • File-backed state: consistent gains (75.6 on Live-SWE, 58.3 on OSWorld).
  • Self-evolution: best SWE score (78.8) but token-expensive.
  • Multi-candidate search: agent calls jump 1.1 → 5.7, yet performance *drops* (73.0 → 71.4) — more search ≠ better harnessing.
  • Context compression: no gain at all — compression loses information agents can't verify.
  • Why it matters

    Harness design becomes a scientific, ablatable research object for the first time. The paper positions NLAH as a midpoint on a spectrum from restrictive code harnesses toward self-harnessing, where a model controls itself without an external layer.

    Limitations

  • Higher token consumption and call counts than code implementations (prototype runtime overhead).
  • Low handoff recall in parent-child setups.
  • Natural language can't faithfully represent mechanisms relying on hidden server-side state.
  • Runtime charter may absorb behavior that should be attributed to the NLAH text (attribution pollution).
  • Takeaways

  • Researchers: publish harness policies as readable .md documents; report module ablations and mechanism audits, not just aggregate scores.
  • Engineers: separate policy (belongs in NLAH) from mechanism (belongs in code); extend AGENTS.md / CLAUDE.md / SKILL.md experience to the run-level harness.
  • Framework authors: make LangChain/CrewAI/AutoGen defaults configurable as NLAH documents backed by an IHR-style shared runtime.

Tags

#ai-agents#agent-harness#nlah#llm-research#paper-review#swe-bench#ablation-study#agent-frameworks

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981396