vLLM says prime-rl runs trillion-scale agentic RL on 28 H200 nodes
vllm_project · x · 2026-07-23
vLLM Project says Prime Intellect’s prime-rl 0.6.0 is running trillion-scale agentic RL on vLLM, using FP8, wide expert parallelism, prefill/decode disaggregation, KV cache offloading, and vllm-router. The setup is used to train GLM-5 on SWE tasks with a 131k sequence length and sub-5-minute steps across 28 H200 nodes.
The post also points readers to a live vLLM Office Hours session featuring a deep dive on prime-rl’s performance and a broader update on vLLM 0.25.
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11