Math-Shepherd replaces 800K human labels with rollouts, lifting GSM8K from 77.9% to 84.1%
le_james94 · x · 2026-09-16
In the context of verification becoming the job: process reward models score every reasoning step, and training one took 800,000 human labels. Math-Shepherd swaps the annotator for rollouts — a step is good if the model can still reach the right answer from it — improving GSM8K from 77.9% to 84.1%.
Related event: Sampling Power: Rollouts Slash Labeling and Boost Solve Rates(2 posts)→
More from Research
- AI2's NGU sampling fixes RL for LLMs that only improves easy tasks — allenai · 2026-09-16
- Jev Benchmark Launches: $42 per Billion Input Tokens, Output Free Forever — cephaloform · 2026-09-16
- Multi-agent RL post-training fights LLM mode collapse and boosts response diversity — natashajaques · 2026-09-16
- World Models Won't Get Us to AGI — Continual Learning Is the Missing Piece, and It's Hard — Intelligent-Cream-14 · 2026-09-16
- Philosopher adds appendix arguing we can be confident today's LLMs are not conscious — AnnaCiaunica · 2026-09-16
- Princeton builds an erasable, light-programmed ultrathin semiconductor a few molecules thick — MengdiWang10 · 2026-09-16