Evolution strategies beat GRPO on reasoning diversity and Pass@K, paper finds

Yunpeng Ba · hf · 2026-08-28

"Understanding Evolution Strategies for LLM Reasoning" examines training LLM reasoning with evolution strategies (ES). Through sparse functional updates and population diversity, ES achieves broader reasoning coverage and higher Pass@K than GRPO. The authors propose a hybrid training approach that combines the strengths of both methods.

Original post →

More from Models

Models channel →