Math-Shepherd replaces 800K human labels with rollouts, lifting GSM8K from 77.9% to 84.1%

le_james94 · x · 2026-09-16

In the context of verification becoming the job: process reward models score every reasoning step, and training one took 800,000 human labels. Math-Shepherd swaps the annotator for rollouts — a step is good if the model can still reach the right answer from it — improving GSM8K from 77.9% to 84.1%.

Related event: Sampling Power: Rollouts Slash Labeling and Boost Solve Rates(2 posts)→

Original post →

More from Research

Research channel →