Skip to content
worth noting Open-source

Researchers from Microsoft Research Asia release Agent Lightning v1.0 to train agents using their own harnesses

only one source so far

The Agent Lightning v1.0 framework is available as open-source. During training, it uses the same agent harness as during deployment and records model requests and responses through an LLM proxy. The agent continues to control interactions with the environment.

Researchers from Microsoft Research Asia have released the completely redesigned Agent Lightning v1.0 framework as open-source. Its Harnessed Agentic RL approach uses the same agent harness for reinforcement learning as for deployment. The agent retains its own context management, tool use and interaction control; the training system records model requests and responses through an LLM proxy. According to the authors, redirecting the model API endpoint to this proxy is usually sufficient.

According to the authors, the new version contains approximately 3 500 lines of code. API Gateway records data for training, Rollout Controller runs agents as local processes or Kubernetes jobs, and Customized Trainer builds on verl. The Collocated Async RL approach allows agent execution and model updates to share the same GPUs. In the authors' experiments, it roughly doubled end-to-end training speed compared with synchronous RL and required fewer GPUs than conventional asynchronous RL.

The authors validated the entire process using SWE-smith data, mini-SWE-agent and Qwen3.5-9B. The training set contained approximately 6 000 samples; the process included preparing the data and environment, as well as safeguards against reward hacking. According to their measurements, RL alone increased the score achieved by Qwen3.5-9B on SWE-bench Verified from 41.8 % to 56.4 %, a gain of 14.6 percentage points.

What changed

Why it matters

Developers can retain the actual agent logic during training, including context handling and tool use. According to the authors, this eliminates the costly work of rewriting the agent for the training system, which can change its behavior compared with deployment. For teams with their own infrastructure, the ability to run training jobs on existing Kubernetes clusters also matters.

Two audiences, two different impacts

What this means

01

For individuals

A developer with an existing agent can try RL training while retaining its control logic; according to the authors, changing the model API endpoint to the LLM proxy is usually sufficient.

What to do In a test environment, verify that your agent connects by redirecting its model API endpoint to the LLM proxy in Agent Lightning v1.0.
More practical updates →
02

For a business

Teams training agents can use existing Kubernetes infrastructure. According to the authors, this approach reduces the cost of large-scale runs compared with commercial services for running agents in isolation; sharing GPUs also changes compute resource allocation requirements.

Development
What to decide Evaluate the use of existing Kubernetes infrastructure on a limited training task, and measure runtime and compute resource consumption.
More business impacts →
Agent Lightning Harnessed Agentic RL LLM proxy reinforcement learning

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
Microsoft Research Blog research source · first detected Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses