Nereus: Adaptive Parallelism for LLM Post-Training
Songlin Jiang, Tuo Shi, Sitong Zhang, Zeke Wang, Mario Di Francesco, Bo Zhao
cs.DC, cs.AI
2026-09-28
Nereus adapts RL post-training parallelism online. A real-data trace shows 27.7% lower step latency vs fixed TP/PP; 8B PPO throughput is 2.14–7.27× OpenRLHF and 1.10–1.47× Verl.
RL post-training for LLMs runs an actor together with critic, reward, and reference models across generation, inference, and training. A single job can burn 100,000 GPU-hours. Three things move while it runs: spot GPUs vanish and schedulers reclaim nodes; the actor learns longer trajectories, and on Llama-3.1-8B generated length grows from about 500 tokens to 8,000 within 1,000 steps, a 16× jump; contention then makes the same layout slower. The paper calls this mix of supply, demand, and hardware-efficiency change drift.
Verl, OpenRLHF, and AReaL freeze DP/TP/PP at launch. Elastic trainers usually manage one model. DynaRL reallocates inside a fixed pool with per-component migration. Three questions stay open together: when a switch is worth its cost, which distributed state can stay on GPU, and how to order GPU handoffs when the job already fills the cluster. Provisioning for the final sequence length wastes GPUs. Provisioning for the start OOMs later.
Nereus is a cost-aware runtime on top of OpenRLHF, Verl, and Laminar, wired to vLLM, DeepSpeed, and Megatron-LM.
A monitor tracks sequence length, the GPU pool, peak memory, and compute/communication efficiency. A pool change or a predicted OOM replans at once. Sequence length drifting more than 30% from the last replan also fires. Switches land only at safe boundaries: the end of an RL step when the job is synchronous, or a weight sync when it is asynchronous. In-flight rollouts keep their policy versions.
The planner picks DP/TP/PP and a GPU set for each model-stage (one model bound to one stage) to minimize predicted steady-state step latency under per-GPU memory and per-stage GPU budgets. The cost model starts from peak FLOPs and link bandwidth, then calibrates with live kernel and collective times, keeping 10% memory headroom. An offline predictor picks an infeasible plan at 8 GPUs and can be up to 1.56× slower than the measured best at 64–256 GPUs.
If the current plan is already illegal, the urgent path skips the economic test, may shrink per-replica batch or turn on recomputation, and falls back to checkpoint/restart. Otherwise a switch must repay its transition cost within γ times the estimated interval until the next threshold crossing, with γ=0.5 by default. That rule is payback admission.
An Elastic Model Unit (EMU) is one data-parallel replica of a model-stage: TP and PP live inside the unit, DP is the replica count, and parameters are not ZeRO-sharded across DP. Split breaks a tight TP/PP unit into looser replicas; Merge fuses them; Extend copies onto new GPUs; Destroy drops a redundant replica. Local order is Split, Destroy, Extend, Merge. When two stages want the same GPUs, the engine adds a release-before-acquire edge and runs ready ops concurrently over NCCL/RCCL.
TP/PP ranks have to collective together. DP replicas only sync now and then, so the state boundary sits at the replica. Scaling the same 8B actor/critic from 16 to 32 GPUs takes 836.74 s with UCP checkpoints, 66.43 s with Tenplex shard moves, and 6.52 s with EMUs.
Three clusters: up to 1,024 MI250X, 256 A100 64GB, and 64 H200. The default workload is Llama-3.1-8B PPO; Qwen3-14B/32B, Llama-3.3-70B, ReMax, GRPO, and async RL are also measured. Each baseline keeps the fastest launch-time layout for that GPU budget.
| Baseline | Setting | Nereus |
| OpenRLHF | 8B PPO throughput | 2.14–7.27× (median 3.99×), 7.27× at 64 GPUs |
| Verl | same, NVIDIA clusters | 1.10–1.47× (median 1.21×) |
| Frozen TP/PP + DP scaling | real-trace step latency | 861.9 s vs 1,191.6 s (−27.7%) |
| DynaRL-style admit | three held-out traces | 858.7 vs 928.3 s/step |
| Empirical optimum (18 settings, 480 plans) | step-latency gap | all ≤5%, exact in 61.1% |
On cluster #2 at fixed budgets, step latency drops by up to 86.3% vs OpenRLHF and 31.9% vs Verl. Generation, inference, and training drop by up to 11.5%, 41.6%, and 29.6% vs Verl. Scaling 32 to 1,024 GPUs cuts Nereus step latency 15× against 10× for OpenRLHF, and the 1,024-GPU job is still 3.20× OpenRLHF. Six switches in a 1,000-step run use 49.5 s of 62,353 s (0.079%). The 512→1,024 jump takes 31.55 s versus 1,629 s for UCP.
Training-stage latency MAPE is 4.96%, max error 14.61%. Plan search is 338 ms at 1,024 GPUs versus 511 s for SCIP. At γ=1 the job takes 14 switches and 939.5 s/step, worse than two switches at 861.9 s. ReMax and GRPO cut latency vs Verl by up to 13.3% and 15.8%; async vs Laminar by up to 37.3%. Over 50 steps both hit reward 0.90 at step 32; at step 50 Nereus is 0.9199 vs 0.9160. Resource scaling is 3.8–16.2× faster than Oobleck and Tenplex. Coordinated overlap trials succeed in all 100 runs per type; per-component migration succeeds in 34–62%.
This is a systems change, not a new RL update rule. Teams already on Verl or OpenRLHF who fight growing sequences and sticky GPU pools are the audience.
The 1.10–1.47× over Verl is the increment on a strong baseline. The 7.27× over OpenRLHF partly reflects a slower baseline. Switch cost is noise at 1,024 GPUs. On the first trace, switching every three steps is already fast; Nereus is only 2.5% faster than that, and 8.3% faster than admitting any positive ΔL.
It is not a drop-in product. It binds vLLM, DeepSpeed, and Megatron, stays on ZeRO-0, and searches power-of-two TP/PP. If a team already hand-tunes close to Verl's best layout, the gain is incremental. If they freeze the launch layout while sequences grow 16×, 27.7% is a real number.
There is no Limitations section. The future-work paragraph is the admission: no context or expert parallelism, no autoscaling of external tool services, no extra RL frameworks.
Training behavior is compared for 50 steps. Reward and KL matching there does not carry to a full post-training run. Actor, reference, reward, and critic are always the same size; real PPO often uses a smaller reward model. Parameters are not sharded across DP, so optimizer state replicates with DP, which is a different memory story than ZeRO-3. "Within 5% of optimal" is inside the controller's own 480-plan space. Verl was not run on MI250X, so the 1,024-GPU numbers are versus OpenRLHF only. Cost-model max error is 14.61%, and the paper does not report how often the urgent path falls through to checkpoint/restart.