What Orchard Is
Microsoft Research pushed Orchard to its open-source main repo in late July, with the latest commit adding MSR-Orchard/slime as a submodule to the trainer. Orchard is not an agent framework—it is an environment layer for agent training, addressing a painful engineering problem: "I want to train ten thousand agents running experiments, but every run needs its own isolated sandbox."
The technical core is Orchard Env: a Kubernetes-native environment service plus Python SDK that can launch thousands of isolated containers on demand over HTTP, offering sandbox lifecycle management, command execution, file I/O, network control, and agent integration. The service stays neutral to agent harnesses, training pipelines, and inference backends—the same Env serves SFT trajectory distillation, RL rollouts, and evaluation task switching without rebuilding the whole system per task.
In the arXiv paper's comparison table, Orchard Env is the only open-source solution meeting the "thin, independent, reusable" definition of an environment service. On cost, normalized per 2-vCPU/8-GiB sandbox, it runs at roughly 0.47x of Daytona (0.10x with spot instances).
Three Agent Training Recipes
Orchard ships three training recipes in one release:
- Orchard-SWE: Starting from Qwen3.5-35B-A3B, it performs "credit-assignment SFT" on 107K distilled trajectories (from MiniMax-M2.5 and Qwen3.5-397B)—learning only productive segments of unresolved trajectories—then stacks Balanced Adaptive Rollout for RL. On SWE-bench Verified it goes from a 61.4% baseline to 69.1% (+BAR) and 69.7% (+dense reward), then to ~73% with 4B value-model reranking—approaching closed-source systems 10x its size.
- Orchard-GUI: A 4B vision-language model as a browser agent, trained with only 400 distilled trajectories plus 2,200 open-ended training tasks. Results: WebVoyager 74.1%, Online-Mind2Web 67.0%, DeepShop 64.0%, averaging 68.4%—current SOTA among open-source GUI agents, competitive with OpenAI/Google CUA.
- Orchard-Claw: A 30B-A3B personal assistant agent starting from 200 synthetic tasks. On Claw-Eval it achieves pass@3 = 59.6%, and 73.9% with the ZeroClaw harness. Crucially, training runs in real deployment harnesses (ReACT, ZeroClaw, OpenClaw, Codex) rather than simplified loops—under the Codex harness, performance improves from 18.6% to 51.5%.
- Code: https://github.com/microsoft/Orchard
- Paper: https://arxiv.org/html/2605.15040
OpenForge: Closing the Train-Deploy Gap
A July update to OpenForge RL directly addresses the training-deployment gap: traditional RL training stacks run simplified harnesses, and performance collapses when swapping to real harnesses at deployment. OpenForge adds a lightweight proxy that records inference calls from real deployment harnesses, reconstructs them as RL samples for any RL codebase (veRL, etc.), while Orchard Env spins up matching containers for rollouts.
Results: OpenForge-GUI 8B scores 37.7 on OSWorld-Verified and 72.3 on WebVoyager; OpenForge-Claw 30B-A3B scores 33.7 on QwenClawBench and 28.1 on MCPAtlas. Agent training can now genuinely happen "in the deployment harness" instead of "in a simulator."
Assessment
Orchard's real signal is not the scores but the taxonomy: it treats "environment" as an independent, reusable service outside the training stack. This decoupling lets the same infrastructure serve SWE, GUI, and Claw domains. The implication is that the next phase of agent training is not "bigger models + more trajectories" but "cheaper, more standardized environment pools + cross-harness RL."
For teams building agent training frameworks—whether on LangGraph, ReAct, or in-house stacks—Orchard's design philosophy points to a clear path: make environments a service, not a component; make training a replaceable backend. It mirrors how Cloudflare turned CI/CD into Workflows—the infrastructure layer is being flattened by the same abstraction.