Meta AI’s OC-GRPO adds 13.8% relative gain on Qwen2.5-7B-Instruct

bronzeagepapi · x · 2026-07-23

A Meta AI paper introduces Off-Context GRPO (OC-GRPO), a small modification to GRPO designed for hard RLVR problems where unguided rollouts get zero reward.

Core idea

Key findings

Takeaway

The method suggests that guided RL training only works reliably when the update remains aligned with the unguided objective.

Original post →

More from Research

Research channel →