RL Training Infra Breakdown: Router Replay, Off-Policy Controls, Full-Vocab OPD on 40+ Teachers

nrehiew_ · x · 2026-09-11

nrehiew details the RL infra behind the new model: dispatch strategies eliminating long-tail stalls, router replay from previous checkpoints, dataset-level caps and discard schemes for short completions, off-policy ratio bounding with loss masking, persistent KV caches and routers on checkpoint updates, and full-vocab OPD distillation on over 40 teacher models at the final stage.

Related event: DeepSeek V4.1 Tech Report Deep Dive: KV Compression and Numeric Reasoning Effort Steal the Show(9 posts)→

Original post →

More from Infra

Infra channel →