KAIST proposes EAPO: entropy-guided credit assignment reinforcing surprising success in RLVR

kaist-ai · hf · 2026-09-29

KAIST AI introduces Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method for RLVR training of LLM reasoning. Motivated by the observation that success under uncertainty is less repeatable while confident failures recur, EAPO couples normalized policy entropy with the sign of response advantage.

Design

Results: best overall performance across reasoning tasks on both base and reasoning backbones, with broader problem coverage and more diverse candidate answers.

Original post →

More from Research

Research channel →