vLLM says prime-rl runs trillion-scale agentic RL on 28 H200 nodes

vllm_project · x · 2026-07-23

vLLM Project says Prime Intellect’s prime-rl 0.6.0 is running trillion-scale agentic RL on vLLM, using FP8, wide expert parallelism, prefill/decode disaggregation, KV cache offloading, and vllm-router. The setup is used to train GLM-5 on SWE tasks with a 131k sequence length and sub-5-minute steps across 28 H200 nodes.

The post also points readers to a live vLLM Office Hours session featuring a deep dive on prime-rl’s performance and a broader update on vLLM 0.25.

Original post →

More from Infra

Infra channel →