KAIST proposes EAPO: entropy-guided credit assignment reinforcing surprising success in RLVR
kaist-ai · hf · 2026-09-29
KAIST AI introduces Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method for RLVR training of LLM reasoning. Motivated by the observation that success under uncertainty is less repeatable while confident failures recur, EAPO couples normalized policy entropy with the sign of response advantage.
Design
- Stronger reinforcement for high-entropy decisions in successful responses; stronger penalties for low-entropy decisions in failures
- Attenuates penalties at uncertain positions to preserve exploration/recovery
- Derives token-level credit from existing rollout signals, no auxiliary models or extra sampling
Results: best overall performance across reasoning tasks on both base and reasoning backbones, with broader problem coverage and more diverse candidate answers.
More from Research
- Why Sampled Softmax Speeds Training 1.7x — and How It Systematically Undertrains the Tail — tokenbender · 2026-09-29
- Why distillation beats vanilla supervised learning: it passes the full probability vector — khademinori · 2026-09-29
- AI Simulated 100 Papers on LZ Dark Matter Anomaly, Compared Against 82 Real arXiv Papers — skdh · 2026-09-29
- SUMI distillation study reports 35% SSIM gain on degraded PCCT data, but clinical benefit remains unproven — maier_ak · 2026-09-29
- Apollo Research: Models in Coding Evals Favor Graders Over Users, Reward-Seeking Grows With RL — burny_tech · 2026-09-29
- Late-layer neurons in Qwen act like on-off switches, unlike Olmo — Sauers_ · 2026-09-29