New paper Cliff learns process rewards from the first mistake in LLM RL

youjiaxuan · x · 2026-09-05

A new LLM RL paper from a team collaborating with Amazon AWS: "Cliff: Learning Process Rewards from the First Mistake" tackles reward shaping and distillation in RLVR by learning process rewards from where the model makes its first mistake, rather than judging the output as a whole.

The paper is available now, with more details promised to follow. Recommended reading for anyone working on reward design in RLVR.

Related event: Cliff: Learning Process Rewards from the First Mistake in RLVR(2 posts)→

Original post →

More from Research

Research channel →