A ladder of weaker feedback signals: ReAct, RLEF, Constitutional AI — and Math-Shepherd's rollouts
le_james94 · x · 2026-09-16
This post frames a 'ladder of weaker signals' in AI feedback loops: an environment answers (ReAct), a test suite answers (RLEF), a written document answers (Constitutional AI) — the loop never changes, but each source is cheaper to obtain and vaguer to act on. Above it: process reward models score every reasoning step (training one took 800K human labels), while Math-Shepherd swaps annotators for rollouts — a step is good if the model can still reach the right answer from it — lifting GSM8K from 77.9% to 84.1%.
Related event: Three Weeks Through Stanford CS329A: Generators Have Outrun Verifiers(9 posts)→
More from Research
- ARCADIA integrates single-cell RNA-seq and spatial proteomics in Bioinformatics — elhamazizi · 2026-09-16
- Minsky's 1960s robot arm solved the same class of problems as today's foundation models — GlenBerseth · 2026-09-16
- Humans scale superlinearly on 14-day tasks while agents hit log-linear limits — AlexGDimakis · 2026-09-16
- Old World journal launches third best-article competition for AI-driven humanities research — ArtificialOther · 2026-09-16
- Hugging Face releases a beginner-friendly visual guide to flow matching — Hugging Face · 2026-09-16
- GPT Astra improves Anthropic's zeta zeroes certificate from 67.25% to 67.31% — dsanft · 2026-09-16