HarnessX Source-Level Architecture Deep Dive
This is an English summary of a Chinese technical post by author "Xiaokai," based on deep study of the open-source HarnessX code by the Darwin Agent team.
- Source: https://github.com/Darwin-Agent/HarnessX
- License: MIT
- Event-driven pipeline: Agent lifecycle is split into 8 hook points:
task_start→step_start→before_model→after_model→before_tool→after_tool→step_end→task_end. Each hook runs a chain of event-stream Processors. - Processor-as-plugin: A minimal Protocol —
async def process(self, event: Event) -> AsyncIterator[Event]— supports pass-through (yield event), transformation (yield modified_event), interception (yield nothing), splitting (multiple yields, e.g., spawning sub-agents), and aborting (raise). - Immutable composition:
HarnessBuilderuses a fluent API (builder.add(proc).slot(...).add_tool(...)) where every mutation returns a new instance; builders merge viabuilder_a | builder_bwith conflict detection (slots, tool names, singleton groups). step_start: system prompt must stay at position 0 unmodified.before_model(strictest): messages cannot be empty; length delta must be 0 or +1; appending is only allowed as a user message when the last message is not already user; only one +1 append is allowed across the whole chain.- Other hooks disallow any message length changes.
- Enforced via
HARNESSX_CONTRACT_MODE:warn(default) orstrict(throwsContractViolationError) — a two-phase rollout strategy. - Phase 0 (TaskStart): system prompt assembly; Phase 1 (StepStart): context assembly via processors; Phase 2 (BeforeModel): final modifications plus safe defaults; Phase 3: model call; Phase 4 (AfterModel): sub-agent spawn, token/cost tracking; Phase 5: tool execution with interrupts (
interrupt_on), before/after-tool processors, turn budget with disk spill; Phase 6 (StepEnd): trajectory recording; Phase 7: termination checks (finish reason, continuation injection, budget); Phase 8 (TaskEnd): evaluation. - Features interrupt/resume via
exit_reason='interrupted'andinterrupted_at,SegmentBoundaryEventfor journal rotation (each segment independently recoverable), and a dual-track message system: append-onlyraw_messagesvs. effectivemessages, with the invariantlen(raw_messages) == len(messages). Stateincludes slots (namespaced dynamic key-value store with Hydra-style_target_serialization), budgets (max_steps, token, cost), and sub-agent tracking.TrajectorySteprecords a full state snapshot (z_t), a state delta (Δz_t, O(changes) memory), the model action, tool observations, and a reward — designed as the core data structure for RL training.- Context:
SystemPromptProcessor(default/Jinja2/null builders),UserWrapperProcessor(XML, CoT),EnvironmentContextInjector. - Control/Safety:
LoopDetectionProcessor(exact-match fingerprints: warn at 3, abort at 5 repeats; compaction-aware cache reset),CompactionProcessor,ToolCallCorrectionLayer,ParseRetryProcessor,SelfVerifyProcessor,TodoWriteEnforcer,CostGuardProcessor(warns at 70% budget),TokenBudgetProcessor,ToolFailureGuard,RepeatedFileEditDetector,BgInstallGuard,SycophancyDetector. - Evaluation:
EvaluationProcessorwithLLMJudgeEvaluator,SelfVerifyEvaluator, and PRM variants (TerminalPRM, DiscountedPRM, ToolSuccessPRM, LLMJudgePRM). - Memory:
MemoryExtractionProcessor(oldest-messages policy) andMemoryRetrievalProcessor(sliding window, summarization, custom backends). - Multi-model:
ModelRouterProcessorwith provider groups and fallback. - Observability:
OTelProcessor(OpenTelemetry),CheckpointProcessor,EpisodeMetricsProcessor. - Tools:
ProgressiveSkillLoader(keyword-matched SKILL.md injection),ToolFilter/ToolWhitelist,ModelSchemaAdapter. - YAML config with Hydra-style
_target_or shorttypenames; plugin system exposing processors, tools, slash commands, MCP servers, lifecycle hooks, and skill directories. HarnessXAgentLoop(registered asharnessx_agent) runs multi-turn rollouts with a state machine (PENDING → GENERATING → PROCESSING_TOOLS → TERMINATED), concurrent tool execution via semaphore, and automatic response truncation.- Example reward:
0.8*accuracy + 0.1*format + 0.1*tool_call; trajectories feed PPO/GRPO policy-gradient updates. - AST semantic hashing of processor code (docstrings/comments stripped) to detect version drift during config serialization.
- Topological ordering: processors sort by
_orderthen_afterdependencies (Kahn's algorithm; cycles raiseHarnessConflictError). - Error isolation: control-flow exceptions (
HarnessError,ContractViolationError) propagate; other processor crashes are logged and events pass through, so one buggy processor cannot kill the run. - HarnessX source: https://github.com/Darwin-Agent/HarnessX
- veRL: https://github.com/volcengine/verl