GEPO adds group-level entropy control to GRPO and wins across 13 benchmarks
internlm · hf · 2026-07-21
What it does
GEPO is a lightweight extension to GRPO for RL training of LLMs that accounts for group-level entropy heterogeneity across mixed tasks.
Key idea
- Global or token-level entropy control can miss the fact that different prompt groups naturally sit in different entropy regimes.
- GEPO estimates entropy from existing grouped samples and applies entropy-conditioned asymmetric advantage shaping.
- It down-weights positive advantages in low-entropy groups to avoid over-exploitation, and preserves negative advantages in high-entropy groups to keep exploration alive.
- Adaptive thresholds are derived from historical entropy statistics.
Results
Across two base models and 13 benchmarks in math, physics, science, code generation, and instruction following, GEPO consistently outperforms GRPO and recent entropy-controlled methods while keeping task-specific exploration balanced.
More from Research
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11