Sampling Power: Rollouts Slash Labeling and Boost Solve Rates
Math-Shepherd replaced 800K human labels with rollouts, lifting GSM8K to 84.1%, while DeepSeek-Coder jumped from 15.9% to 56% on SWE-bench Lite with 250 samples, showing generators far exceed single-attempt performance.
2026-09-16 ~ 2026-09-16 · 2 related posts
- Repeated sampling lifts DeepSeek-Coder on SWE-bench Lite from 15.9% to 56% — le_james94 · 2026-09-16
- Math-Shepherd replaces 800K human labels with rollouts, lifting GSM8K from 77.9% to 84.1% — le_james94 · 2026-09-16