Microsoft Research Asia has open-sourced Agent Lightning v1.0, a lightweight framework implementing the "Harnessed Agentic RL" training paradigm. This approach allows reinforcement learning to use the same agent harness deployed in production, eliminating the need to reimplement agent logic within the training system.

  • The framework consists of approximately 3,500 lines of code and includes an API Gateway, Rollout Controller, and Customized Trainer.
  • It supports native Kubernetes execution for agents on self-managed clusters, cloud infrastructure, or local environments without commercial sandbox dependencies.
  • An end-to-end coding agent pipeline using Qwen3.5-9B achieved a 14.6 percentage point gain on SWE-bench Verified, raising Pass@1 from 41.8% to 56.4% with only about 6,000 training samples.

This design enables researchers and developers to train agents directly with their existing deployment harnesses, ensuring the trained model behaves identically to the deployed system while maintaining a simple, reproducible pipeline.