Researchers from Microsoft Research Asia release Agent Lightning v1.0 to train agents using their own harnesses
The Agent Lightning v1.0 framework is available as open-source. During training, it uses the same agent harness as during deployment and records model requests and responses through an LLM proxy. The agent continues to control interactions with the environment.
Researchers from Microsoft Research Asia have released the completely redesigned Agent Lightning v1.0 framework as open-source. Its Harnessed Agentic RL approach uses the same agent harness for reinforcement learning as for deployment. The agent retains its own context management, tool use and interaction control; the training system records model requests and responses through an LLM proxy. According to the authors, redirecting the model API endpoint to this proxy is usually sufficient.
According to the authors, the new version contains approximately 3 500 lines of code. API Gateway records data for training, Rollout Controller runs agents as local processes or Kubernetes jobs, and Customized Trainer builds on verl. The Collocated Async RL approach allows agent execution and model updates to share the same GPUs. In the authors' experiments, it roughly doubled end-to-end training speed compared with synchronous RL and required fewer GPUs than conventional asynchronous RL.
The authors validated the entire process using SWE-smith data, mini-SWE-agent and Qwen3.5-9B. The training set contained approximately 6 000 samples; the process included preparing the data and environment, as well as safeguards against reward hacking. According to their measurements, RL alone increased the score achieved by Qwen3.5-9B on SWE-bench Verified from 41.8 % to 56.4 %, a gain of 14.6 percentage points.
Why it matters
Developers can retain the actual agent logic during training, including context handling and tool use. According to the authors, this eliminates the costly work of rewriting the agent for the training system, which can change its behavior compared with deployment. For teams with their own infrastructure, the ability to run training jobs on existing Kubernetes clusters also matters.
Two audiences, two different impacts
What this means
For individuals
A developer with an existing agent can try RL training while retaining its control logic; according to the authors, changing the model API endpoint to the LLM proxy is usually sufficient.
For a business
Teams training agents can use existing Kubernetes infrastructure. According to the authors, this approach reduces the cost of large-scale runs compared with commercial services for running agents in isolation; sharing GPUs also changes compute resource allocation requirements.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.