Evolution Strategies Boost LLM Reasoning Coverage

yeewhye · x · 2026-09-01

Research indicates that Evolution Strategies (ES) can post-train LLMs without backpropagation. Compared to GRPO, ES explores higher Pass@K, achieves sparse functional updates without catastrophic forgetting, and requires fewer samples as models scale. ES complements GRPO: ES offers broader reasoning coverage while GRPO optimizes Pass@1; sequential combinations yield better Pass@1–Pass@K trade-offs.

Original post →

More from Research

Research channel →