Инструменты

Microsoft представила Agent Lightning: легковесный фреймворк для обучения ИИ-агентов

Исследователи из Microsoft Research Asia выпустили Agent Lightning v1.0 — систему из 3500 строк кода, позволяющую обучать агентов с помощью обучения с подкреплением без переписывания их архитектуры.

Обновлено:
4 мин чтения
Microsoft представила Agent Lightning: легковесный фреймворк для обучения ИИ-агентов

Суть

Исследователи из Microsoft Research Asia представили Agent Lightning v1.0 — новую парадигму и легковесный фреймворк для обучения ИИ-агентов с помощью обучения с подкреплением (RL). Главное нововведение заключается в том, что система позволяет использовать ту же самую программную обвязку (harness), которая применяется при развертывании агента, напрямую в процессе обучения. Это избавляет разработчиков от необходимости заново реализовывать логику агента внутри обучающего фреймворка.

Контекст

Figure 1: Side-by-side comparison of two training loops. In Agentic RL, the environment exchanges actions and observations with a tokenizer, which passes action and observation tokens to the policy model. In Harnessed Agentic RL, an agent harness handling context and orchestration sits between the environment and an OpenAI-like API, which exchanges input and output tokens with the policy model.

Figure 1: Side-by-side comparison of two training loops. In Agentic RL, the environment exchanges actions and observations with a tokenizer, which passes action and observation tokens to the policy model. In Harnessed Agentic RL, an agent harness handling context and orchestration sits between the environment and an OpenAI-like API, which exchanges input and output tokens with the policy model.

Современные ИИ-агенты эволюционировали от одиночных моделей до сложных систем, включающих инструменты и среды выполнения. Их возможности во многом зависят от внешней обвязки, координирующей работу. Обучение с подкреплением помогает улучшать такие системы методом проб и ошибок. Однако традиционные системы (например, verl, AReaL и slime) предполагают, что обучающий фреймворк полностью управляет циклом взаимодействия со средой. Из-за этого разработчикам приходилось переписывать код агентов для обучения, что дорого и приводит к тому, что обучаемый агент отличается от того, который в итоге будет развернут.

Детали

Agent Lightning v1.0 решает эту проблему, помещая прокси-сервер большой языковой модели (LLM proxy) между агентом и моделью. Агент продолжает работать как обычно, а фреймворк наблюдает и записывает его вызовы.

Ключевые характеристики системы:

  • Компактность: весь фреймворк состоит примерно из 3500 строк кода и включает три основных компонента: API-шлюз, контроллер развертывания и настроенный тренер.
  • Встроенная поддержка Kubernetes: агенты запускаются как стандартные задачи Kubernetes на собственных кластерах или локальной инфраструктуре, что исключает зависимость от платных коммерческих песочниц (таких как Modal Sandbox или E2B).
  • Эффективность данных: в тесте с кодовым агентом пайплайн позволил повысить показатель Pass@1 модели Qwen3.5-9B на SWE-bench Verified с 41.8% до 56.4% (прирост на 14.6 процентных пункта) с использованием около 6000 обучающих примеров.
  • Collocated Async RL: новый подход, при котором сбор данных (rollout) и обновление модели делят одни и те же графические процессоры (GPU). В экспериментах это дало примерно двукратное ускорение по сравнению с синхронным RL при меньшем количестве используемых GPU.
Figure 2: System architecture diagram. On the left, agents with harnesses — mini-SWE-agent, OpenHands, and OpenClaw — run on a Kubernetes cluster. They connect to three components: an API Gateway containing a Rollout API and an LLM API Proxy, a Rollout Controller containing a local reconciler and a Kubernetes reconciler, and a Customized Trainer containing a sample adapter and monitoring. These connect in turn to an inference engine and a training engine holding the model.

Figure 2: System architecture diagram. On the left, agents with harnesses — mini-SWE-agent, OpenHands, and OpenClaw — run on a Kubernetes cluster. They connect to three components: an API Gateway containing a Rollout API and an LLM API Proxy, a Rollout Controller containing a local reconciler and a Kubernetes reconciler, and a Customized Trainer containing a sample adapter and monitoring. These connect in turn to an inference engine and a training engine holding the model.

Анализ

Переход к парадигме Harnessed Agentic RL означает, что среда и цикл взаимодействия теперь управляются самой обвязкой агента, а не обучающей системой. Это создает новые технические вызовы, такие как правильная ретокенизация, расчет преимуществ на уровне сессий (rollouts) и нормализация потерь. Решение этих проблем внутри компактной кодовой базы делает процесс обучения более стабильным. Использование стандартных инструментов вроде Kubernetes снижает вычислительные затраты и барьер входа для масштабирования.

Перспектива

Упрощение интеграции существующих агентов с системами обучения с подкреплением может ускорить развитие автономных систем. Отсутствие необходимости переписывать код для обучения и независимость от проприетарных сервисов делают разработку более воспроизводимой и доступной для исследователей.

Figure 3: Three GPU scheduling timelines. Synchronous RL uses four GPUs at low efficiency, with long idle gaps before a single update block. Collocated Async RL uses the same four GPUs at high efficiency, interleaving full and partial rollouts with update blocks. Asynchronous RL reaches high efficiency but requires eight GPUs. Bars are colored for full rollout, partial rollout, and update.

Figure 3: Three GPU scheduling timelines. Synchronous RL uses four GPUs at low efficiency, with long idle gaps before a single update block. Collocated Async RL uses the same four GPUs at high efficiency, interleaving full and partial rollouts with update blocks. Asynchronous RL reaches high efficiency but requires eight GPUs. Bars are colored for full rollout, partial rollout, and update.

Figure 4: Flow diagram. An API Gateway holds three rollouts, two queueing and one running. The Rollout Controller polls the gateway and uses a Kubernetes reconciler to create jobs on a Kubernetes cluster, and a local reconciler to watch and list local processes. Status updates flow back to the gateway.

Figure 4: Flow diagram. An API Gateway holds three rollouts, two queueing and one running. The Rollout Controller polls the gateway and uses a Kubernetes reconciler to create jobs on a Kubernetes cluster, and a local reconciler to watch and list local processes. Status updates flow back to the gateway.

Figure 5: Two line charts plotting 200 training steps. On the left, validation reward: rollout-level advantage combined with rollout-level normalization reaches the highest reward at about 0.37, above rollout-level advantage alone and sample-level advantage. On the right, policy entropy: rollout-level advantage alone climbs steeply to about 0.65, while the combined method stays lower and steadier.

Figure 5: Two line charts plotting 200 training steps. On the left, validation reward: rollout-level advantage combined with rollout-level normalization reaches the highest reward at about 0.37, above rollout-level advantage alone and sample-level advantage. On the right, policy entropy: rollout-level advantage alone climbs steeply to about 0.65, while the combined method stays lower and steadier.

System architecture diagram. On the left, agents with harnesses — mini-SWE-agent, OpenHands, and OpenClaw — run on a Kubernetes cluster. They connect to three components: an API Gateway containing a Rollout API and an LLM API Proxy, a Rollout Controller containing a local reconciler and a Kubernetes reconciler, and a Customized Trainer containing a sample adapter and monitoring. These connect in turn to an inference engine and a training engine holding the model.

System architecture diagram. On the left, agents with harnesses — mini-SWE-agent, OpenHands, and OpenClaw — run on a Kubernetes cluster. They connect to three components: an API Gateway containing a Rollout API and an LLM API Proxy, a Rollout Controller containing a local reconciler and a Kubernetes reconciler, and a Customized Trainer containing a sample adapter and monitoring. These connect in turn to an inference engine and a training engine holding the model.

Abstract teal-toned image of aluminum cans viewed from above, overlaid with a network of connected flowchart shapes (rectangles, rounded nodes, and diamonds) linked by thin line.

Abstract teal-toned image of aluminum cans viewed from above, overlaid with a network of connected flowchart shapes (rectangles, rounded nodes, and diamonds) linked by thin line.

TL;DR

Главное

Agent Lightning v1.0 позволяет обучать ИИ-агентов с помощью обучения с подкреплением, используя их оригинальный код, что упрощает процесс и снижает затраты на инфраструктуру.

Ключевые факты

  • /Фреймворк состоит всего из 3500 строк кода.
  • /Модель Qwen3.5-9B улучшила результат на SWE-bench Verified с 41.8% до 56.4% (на 14.6 процентных пункта).
  • /Для достижения этого результата потребовалось около 6000 обучающих примеров.
  • /Подход Collocated Async RL обеспечил примерно двукратное ускорение по сравнению с синхронным RL.

Инсайт

Перенос управления циклом взаимодействия от обучающего фреймворка к самому агенту решает проблему расхождения между поведением модели во время тренировки и в реальных условиях развертывания.

Источник:Microsoft

Читайте также