GEPO adds group-level entropy control to GRPO and wins across 13 benchmarks

internlm · hf · 2026-07-21

What it does

GEPO is a lightweight extension to GRPO for RL training of LLMs that accounts for group-level entropy heterogeneity across mixed tasks.

Key idea

Results

Across two base models and 13 benchmarks in math, physics, science, code generation, and instruction following, GEPO consistently outperforms GRPO and recent entropy-controlled methods while keeping task-specific exploration balanced.

Original post →

More from Research

Research channel →