Evolution strategies rival policy gradients for LLM fine-tuning: Best Paper Runner-up

caglarml · x · 2026-10-10

The paper "Dimension Is Free, Horizon Is Not: When Evolution Strategies Rival Policy Gradients for LLM Fine-Tuning" won a Best Paper Runner-up. It shows evolution strategies can rival policy gradients for LLM fine-tuning — dimensions are free, but the horizon is not — offering an alternative to the mainstream policy-gradient RLHF recipe. Authors include Catherine Liang.

Original post →

More from Models

Models channel →