Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Memory
yeewhye · x · 2026-08-20
A new paper proposes Agentic ESOpt, a framework designed to address fine-tuning challenges for long-horizon LLM agents. Unlike traditional Reinforcement Learning (RL), this method uses Evolution Strategies (ES) to optimize in the parameter space. Its advantages include requiring only inference-level GPU memory for full-parameter fine-tuning, supporting co-evolution of prompts and parameters, and better attribution for long trajectories without the complex credit assignment problems found in RL.
Related event: Agentic ESOpt Tunes Long-Horizon Agents with Evolution Strategies(3 posts)→
More from coding & agent
- AI engineering is more like lawmaking than board games, argues Drew Breunig — dbreunig · 2026-09-23
- Seroter's daily digest: GPT-6 and Opus 5.5 ship, 1 in 4 agents run unmonitored — rseroter · 2026-09-23
- Spawning Claude agents that auto-open terminal panes: 'tmux can't do this' — letandrewcook · 2026-09-23
- Toddler's interactive storybook built in two hours with an agent workflow — mimi10v3 · 2026-09-23
- Your data stack is about to get less forgiving: agents need data that's true now — bigdata · 2026-09-23
- Opinion: Agents make software good at using software, not just being used — r0ck3t23 · 2026-09-23