GEPO adds group-level entropy control to GRPO and wins across 13 benchmarks
internlm · hf · 2026-07-21
What it does
GEPO is a lightweight extension to GRPO for RL training of LLMs that accounts for group-level entropy heterogeneity across mixed tasks.
Key idea
- Global or token-level entropy control can miss the fact that different prompt groups naturally sit in different entropy regimes.
- GEPO estimates entropy from existing grouped samples and applies entropy-conditioned asymmetric advantage shaping.
- It down-weights positive advantages in low-entropy groups to avoid over-exploitation, and preserves negative advantages in high-entropy groups to keep exploration alive.
- Adaptive thresholds are derived from historical entropy statistics.
Results
Across two base models and 13 benchmarks in math, physics, science, code generation, and instruction following, GEPO consistently outperforms GRPO and recent entropy-controlled methods while keeping task-specific exploration balanced.
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11