Cliff: learning process rewards from the first reasoning mistake

Peixuan Han · hf · 2026-09-03

Cliff improves RL with verifiable rewards by using an off-the-shelf LLM to detect the first reasoning error and shape token-level advantages accordingly.

Related event: Cliff: Learning Process Rewards from the First Mistake in RLVR(2 posts)→

Original post →

More from Research

Research channel →