vLLM says prime-rl runs trillion-scale agentic RL on 28 H200 nodes
vllm_project · x · 2026-07-23
vLLM Project says Prime Intellect’s prime-rl 0.6.0 is running trillion-scale agentic RL on vLLM, using FP8, wide expert parallelism, prefill/decode disaggregation, KV cache offloading, and vllm-router. The setup is used to train GLM-5 on SWE tasks with a 131k sequence length and sub-5-minute steps across 28 H200 nodes.
The post also points readers to a live vLLM Office Hours session featuring a deep dive on prime-rl’s performance and a broader update on vLLM 0.25.
More from Infra
- Waymo’s open data lets an independent study verify its earlier safety analyses — reed · 2026-07-23
- European AI company’s U.S. cloud deal reignites the sovereignty debate — jonerp · 2026-07-23
- Etched hits a $10.3B valuation with chips it says accelerate inference without GPUs — TechCrunch AI · 2026-07-23
- Daily dashboard shows open-model adoption still led by China and Qwen — natolambert · 2026-07-23
- DeepSeek-V4-Pro training on Ascend 910C exposes long-tail kernel bottlenecks — teortaxesTex · 2026-07-23
- Besi China revenue hits €84.2M in Q1 2026, or 46% of sales — zephyr_z9 · 2026-07-23