Microsoft Research released the open-source Orchard framework for training AI agents on Kubernetes
Microsoft Research released the open-source Orchard framework for training and evaluating AI agents on Kubernetes, including training data and the Orchard-SWE, Orchard-GUI and Orchard-Claw recipes. According to the company, Orchard-SWE raised the score on SWE-bench Verified from 61.4 % to 73 %.
The Microsoft Research division released Orchard, an open-source framework for scalable training and evaluation of agentic AI systems. At its core is Orchard Env – a lightweight environment built on Kubernetes that provides reusable isolated components for running and building agents at scale, from training data collection through reinforcement learning rollouts to evaluation. According to Microsoft Research, unlike many existing frameworks, Orchard Env is designed to support different types of agents and tasks without modification – software-engineering agents, web-browsing agents and personal-assistant agents. The Kubernetes foundation enables thousands of isolated components to be created, managed and destroyed in parallel.
According to the company, the framework also addresses the problem of training within sophisticated harnesses such as Claude Code, Codex or OpenClaw, which manage multi-turn reasoning, tool use and connections to external systems. A lightweight proxy tool records model calls directly from the harness as training data, with each rollout running in its own container – allowing the agent to be trained end-to-end directly in the harness in which it will be deployed (OpenClaw, Codex, ZeroClaw and others), instead of in a simplified substitute.
The release includes three domain-specific training recipes – Orchard-SWE, Orchard-GUI and Orchard-Claw – along with the training data and evaluation methods used to create them. Orchard-SWE is built on the Mini-SWE-Agent framework and evaluated on the SWE-bench Verified benchmark. According to the company, 107 000 agent interactions were distilled for training from two open models (MiniMax-M2.5 and Qwen3.5-397B) across a broad sample of GitHub Issues, using supervised fine-tuning with credit assignment that also makes use of partially successful agent attempts. This is followed by reinforcement learning using Balanced Adaptive Rollout, supplemented with denser reward techniques (on-policy distillation and process reward model), and finally reranking candidate solutions using a 4-billion-parameter value model trained on previous rollouts.
According to Microsoft Research, this approach increased the Orchard-SWE score on SWE-bench Verified from a baseline of 61.4 % to 69.1 % (Balanced Adaptive Rollout), 69.7 % (with denser reward techniques) and 73 % after reranking with the value model – which the company describes as a new best result among open-source models of comparable size (approximately 3 billion active parameters), approaching the results of frontier systems more than ten times larger. Orchard-GUI trains a 4-billion-parameter vision-language model as a browser agent with a smaller amount of supervision (400 distilled examples). Details on Orchard-GUI and Orchard-Claw can be found in the source article.
Why it matters
The framework removes the dependence on proprietary infrastructure previously required to train leading agentic systems and openly provides training data and evaluation methods – enabling researchers and smaller teams to build and test their own agents without access to closed pipelines at major labs. According to the company, the ability to train an agent directly in the harness in which it will be deployed (e.g. Claude Code or Codex) also eliminates the mismatch between training and production environments that has until now affected the results of research tools.
Two audiences, two different impacts
What this means
For individuals
Developers and researchers working with AI agents now have access to an open-source tool, training data and evaluation methods that previously required access to proprietary infrastructure at major labs.
For a business
Companies developing agentic AI systems gain an open-source alternative to proprietary sandboxes and training pipelines; according to Microsoft Research, the framework enables agents to be trained directly within production harnesses (e.g. Claude Code, Codex), eliminating the mismatch between training and deployment environments.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.