Paper Overview
Field: NLP Authors: Haoran Ye, Yuxing Lu, Haonan Dong, Zhaochen Su, Guojie Song Published: 2026-09-21 arXiv: 2609.24974
Abstract
Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones. The authors therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time guidance and transferring the behaviors it induces into model weights, so that its gains survive under a single fixed target harness.
The challenge is that the two harnesses differ in action space and available information, so guidance from the optimized harness cannot serve directly as supervision for the target harness.
Method: Harness-Zero
Harness-Zero performs distillation via Agent-as-Harness:
- Guided by the optimized harness, a *harnessing agent* corrects the student model's responses within the target harness's action space before execution.
- This converts harness guidance into training demonstrations.
- Fine-tuning on the resulting trajectories internalizes harness-induced behaviors into model capabilities, allowing the specialized harness to be removed at deployment.
Results
Experiments spanning knowledge work, tool use, and science domains show:
1. For frontier LLMs using the same evolutionary harness, agent-as-harness outperforms code-as-harness. 2. With the specialized harness removed at deployment, Harness-Zero lifts the base model's macro-average task success rate from 23.3% to 44.3%, even surpassing the 41.7% achieved when the harness was still present. 3. Harness-Zero recovers missing harness-induced behaviors in the base model, restoring on average 82.3% across 28 patterns in three domains.
---
Source: arXiv:2609.24974