Nereus: adaptive parallelism boosts 8B PPO throughput up to 7.27x over OpenRLHF
Songlin Jiang · hf · 2026-09-29
Nereus is a cost-aware runtime that adapts parallelism for LLM RL post-training, where resource availability, sequence lengths, memory pressure, and stage bottlenecks shift during a run, making initial execution plans stale or infeasible.
Design
- A low-overhead controller picks a memory-feasible global plan and uses a cost model calibrated against the running job to decide if a transition is worth it
- Each model-stage replica's distributed state is an Elastic Model Unit; a global transition graph orders transformations and GPU transfers
Results
- Online TP/PP adaptation cuts average step latency 27.7% vs. a fixed TP/PP layout with DP scaling
- Six transitions cost just 0.079% of a 1,000-step run on 1,024 GPUs
- End-to-end 8B PPO throughput: 2.14–7.27x over OpenRLHF, 1.10–1.47x over Verl
More from Infra
- Celesto: open-source persistent microVM computers for AI agents, boots in 500ms — aniketmaurya · 2026-09-29
- Redditor runs gpt-oss-120b across a phone, three Macs and two Windows PCs — ANR2ME · 2026-09-29
- BioNeMo team boosts Mixtral-8x7B training throughput 2.21x vs HF BF16 baseline — AllThingsApx · 2026-09-29
- DeepSeek's elastic compute team is hiring heavily, shares sandbox infra for large-scale agent training — teortaxesTex · 2026-09-29
- Developer slams third-party inference providers: Gemini up 10x, Luna 15s latency — julianharris · 2026-09-29
- Data center water use isn't about total volume, it's who runs out of local freshwater first — AryHHAry · 2026-09-29