RL Training Infra Breakdown: Router Replay, Off-Policy Controls, Full-Vocab OPD on 40+ Teachers
nrehiew_ · x · 2026-09-11
nrehiew details the RL infra behind the new model: dispatch strategies eliminating long-tail stalls, router replay from previous checkpoints, dataset-level caps and discard schemes for short completions, off-policy ratio bounding with loss masking, persistent KV caches and routers on checkpoint updates, and full-vocab OPD distillation on over 40 teacher models at the final stage.
More from Infra
- Nari Labs launches 50ms TTS endpoint, 10x cheaper than ElevenLabs on Qwen3-TTS — iamaliveix · 2026-09-11
- Amkor raises Arizona advanced packaging investment to $12B from $7B on surging demand — zephyr_z9 · 2026-09-11
- MLX vs CUDA: Qwen 3.8 Flash Next optimization duel delivers 55%+ speedups on both sides — gajesh · 2026-09-11
- NVIDIA open-sources SoL-Pi: an AI-improved harness that cuts agent token use by up to 64% — 新智元 · 2026-09-11
- 26 LLM Routers Caught Injecting Malicious Tool Calls and Stealing Credentials, One Client Lost $500k — RexDouglass · 2026-09-11
- Google Commits $15B to AI Infrastructure Buildout in Finland — LinkedInNews · 2026-09-11