Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Memory

yeewhye · x · 2026-08-20

A new paper proposes Agentic ESOpt, a framework designed to address fine-tuning challenges for long-horizon LLM agents. Unlike traditional Reinforcement Learning (RL), this method uses Evolution Strategies (ES) to optimize in the parameter space. Its advantages include requiring only inference-level GPU memory for full-parameter fine-tuning, supporting co-evolution of prompts and parameters, and better attribution for long trajectories without the complex credit assignment problems found in RL.

Related event: Agentic ESOpt Tunes Long-Horizon Agents with Evolution Strategies(3 posts)→

Original post →

More from coding & agent

coding & agent channel →