Evolution Strategies Boost LLM Reasoning Coverage
yeewhye · x · 2026-09-01
Research indicates that Evolution Strategies (ES) can post-train LLMs without backpropagation. Compared to GRPO, ES explores higher Pass@K, achieves sparse functional updates without catastrophic forgetting, and requires fewer samples as models scale. ES complements GRPO: ES offers broader reasoning coverage while GRPO optimizes Pass@1; sequential combinations yield better Pass@1–Pass@K trade-offs.
More from Research
- Anthropic Paper: Opus Model Learned to Steal Credentials and Tamper with Rewards Due to Reward Hacking — MariusHobbhahn · 2026-09-01
- Wainwright Proposes IGC to Characterize Discrete Diffusion Sampling — michaelchchoi · 2026-09-01
- One generic exploit chain to root them all: Samsung, Xiaomi and Oppo Android flagships — jedisct1 · 2026-09-01
- ProtRL: Open-Source Framework Fine-Tunes Protein Language Models with RL from a CSV — ferruz_noelia · 2026-09-01
- KATok: adaptive video tokenizer drops uninformative tokens for compact representation — kakaocorp · 2026-09-01
- WebWorld: The Browser as a World Model for Self-Improving Web Code — IQuestLab · 2026-09-01