RL run spends a lot of compute on graders; is re-prefilling worth it over 30 steps

stochasticchasm · x · 2026-09-22

Noting that a recent RL training run spends comparatively large compute on graders, the author wonders which model was used as grader, and argues that with only 30 total steps the policy shifts significantly each step, so re-prefilling the grader may be worthwhile.

Related event: RL training cost debate: grader compute and KV cache trade-offs(2 posts)→

Original post →

More from Research

Research channel →