GEPO adds group-level entropy control to GRPO and wins across 13 benchmarks
internlm · hf · 2026-07-21
What it does
GEPO is a lightweight extension to GRPO for RL training of LLMs that accounts for group-level entropy heterogeneity across mixed tasks.
Key idea
- Global or token-level entropy control can miss the fact that different prompt groups naturally sit in different entropy regimes.
- GEPO estimates entropy from existing grouped samples and applies entropy-conditioned asymmetric advantage shaping.
- It down-weights positive advantages in low-entropy groups to avoid over-exploitation, and preserves negative advantages in high-entropy groups to keep exploration alive.
- Adaptive thresholds are derived from historical entropy statistics.
Results
Across two base models and 13 benchmarks in math, physics, science, code generation, and instruction following, GEPO consistently outperforms GRPO and recent entropy-controlled methods while keeping task-specific exploration balanced.
More from Research
- Anthropic masterclass spotlights how to build and observe AI agents — _jaydeepkarale · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21