Microsoft’s Agent Lightning Trains AI Agents Inside the Harnesses They Actually Use

October 10, 2026

A working AI agent harness sends real tool and environment interactions through a compact training control plane into a reinforcement-learning model loop.
Agent Lightning keeps the deployed harness intact while its gateway, rollout controller, and trainer convert real agent work into reinforcement-learning updates.

Microsoft Research has presented Agent Lightning v1.0, an open-source framework for improving AI agents with reinforcement learning while preserving the tools, context, control flow, and execution environments used in production.

Its roughly 3,500-line core control plane sits between an existing agent and its model endpoint, capturing the prompts, responses, log probabilities, and rollout events required for training. Developers point the harness at Agent Lightning’s OpenAI-compatible proxy instead of rebuilding the agent inside a separate reinforcement-learning framework.

The system has three main components: an API gateway that records model traffic, a rollout controller that launches agents as local processes or standard Kubernetes jobs, and a customized trainer built on verl. Microsoft’s collocated asynchronous design lets rollouts and model updates share the same GPU pool, pausing new requests only while an update is applied.

Microsoft also published a reproducible coding-agent recipe. In its experiment, training Qwen3.5-9B with approximately 6,000 examples increased SWE-bench Verified Pass@1 from 41.8% to 56.4%—an absolute gain of 14.6 percentage points. This is a Microsoft-reported result from one training configuration, not a guarantee for other models, harnesses, or workloads.

Why it matters

Agent performance depends on more than the underlying model: the harness determines how tools, memory, environments, context management, subagents, and feedback loops work. Agent Lightning offers builders a practical way to train the complete deployed system instead of a simplified imitation of it.

That complements the infrastructure shift behind OpenAI’s managed Agents API and Codex harness. Managed runtimes make full agent systems easier to deploy; Agent Lightning targets the next layer—improving those systems from their real execution traces without replacing their operating logic.

The framework, training pipeline, and scripts are available under the MIT license, which could make harness-aware reinforcement learning more accessible to teams building coding agents, search agents, and other tool-using systems. Teams still need careful reward design, environment isolation, evaluation splits, and safeguards against reward hacking before treating benchmark gains as production improvements.

Relevant links

← Back to stories