Miles ships Day-0 RL support for DeepSeek-V4.1-Flash with KL held at 0.0012–0.0017
ying11231 · x · 2026-09-10
radixark's RL training framework Miles added Day-0 support for DeepSeek-V4.1-Flash:
- Trainer stays close to SGLang sampling; parallelism and shared state let the new architecture scale intact across GPUs
- Quantization-aware training mirrors SGLang's FP4/FP8 rounding, with routing replay reusing rollout expert choices
- FP32 and deterministic reductions ensure numerical consistency; colocated training/rollout fits full-parameter RL on 16 GPUs
Over steps 0–80 of a DAPO run, per-token trainer–rollout KL stayed at 0.0012–0.0017 while reward rose from 0.51 to 0.78. Blog and cookbook in the comments.
More from Infra
- Matt Barrie burned 4B tokens in a day, cut his bill 500-fold, and now worries about $5T in debt — gaganghotra_ · 2026-09-10
- Analyst: DeepSeek's latest change is a big win for token efficiency, moving toward OpenAI's regime — teortaxesTex · 2026-09-10
- Acellera tests 7 LLM+harness combos on drug discovery: one RTX 5090 holds up — gdefabritiis · 2026-09-10
- Mac mini tested: local 35B runtime hits Haiku-level scores but falls short for agents — PawelHuryn · 2026-09-10
- The data center is a symbol: why debunked claims about AI infrastructure still spread — ShakeelHashim · 2026-09-10
- Kimi K3 lands on RunPod: 2.8T params, 1M context, $3/$15 per 1M tokens — Kimi_Moonshot · 2026-09-10