Agon trains reasoning by having models grade each other

Vladislav Beliaev · hf · 2026-07-20

Agon: rival-grading RL for reasoning

The paper argues that verifiable-reward RL methods like GRPO only grade the final answer, which can push models to write more without improving thinking quality. To address that, it introduces Agon, a competitive cross-model RL setup where two models act as each other's graders.

How it works

Why it matters

Because both models are being optimized, each faces a progressively stronger rival, something single-model RL cannot provide. At inference time, the pair is used as a two-stage cascade: one model drafts, the other answers after reading the draft.

Reported results

The authors say the next step is to let the models reason together in latent space instead of text.

Original post →

More from Research

Research channel →