A ladder of weaker feedback signals: ReAct, RLEF, Constitutional AI — and Math-Shepherd's rollouts

le_james94 · x · 2026-09-16

This post frames a 'ladder of weaker signals' in AI feedback loops: an environment answers (ReAct), a test suite answers (RLEF), a written document answers (Constitutional AI) — the loop never changes, but each source is cheaper to obtain and vaguer to act on. Above it: process reward models score every reasoning step (training one took 800K human labels), while Math-Shepherd swaps annotators for rollouts — a step is good if the model can still reach the right answer from it — lifting GSM8K from 77.9% to 84.1%.

Related event: Three Weeks Through Stanford CS329A: Generators Have Outrun Verifiers(9 posts)→

Original post →

More from Research

Research channel →