Long-context RL infrastructure tackles 1M-token training with partial rollouts

SeunghyunSEO7 · x · 2026-07-28

The post points to a long-context RL infrastructure setup for agentic training, including partial rollouts, external KV-cache storage, and resumable execution.

The image shows a paper section on “Infra for 1M Agentic RL,” describing how co-located RL training and partial rollouts help keep 1M-context Kimi K3 experiments within a few hundred GPUs. The tradeoff is heavier memory pressure because rollout KV-cache must be preserved for the next iteration, alongside the training state.

Original post →

More from Infra

Infra channel →