RL run spends a lot of compute on graders; is re-prefilling worth it over 30 steps
stochasticchasm · x · 2026-09-22
Noting that a recent RL training run spends comparatively large compute on graders, the author wonders which model was used as grader, and argues that with only 30 total steps the policy shifts significantly each step, so re-prefilling the grader may be worthwhile.
Related event: RL training cost debate: grader compute and KV cache trade-offs(2 posts)→
More from Research
- Question's Gambit lifts deep research agents: GPT-5.5 hits 90.5% on BrowseComp-Plus — omarsar0 · 2026-09-22
- PyroDash cuts inference cost 96% by having a 4B model call the big one only when needed — jiqizhixin · 2026-09-22
- phantom-kv: uncensor LLMs per-request with an 18MB trained KV-cache, no weight edits — Anony6666 · 2026-09-22
- Lean vs ZFC: the rules of mathematical proof weren't changed by any vote — jessi_cata · 2026-09-22
- David Krueger: four unresolved foundational problems stand between us and safe AI — DavidSKrueger · 2026-09-22
- Full-Parameter RL on TPUs: peano_ai Runs 310B MiMo-V2.6 Across 1,000+ TPUs — simonguozirui · 2026-09-22