Nereus: Adaptive Parallelism for LLM Post-Training
This paper presents Nereus, a cost-aware runtime that adapts LLM reinforcement learning (RL) post-training jobs into efficient execution plans on GPU clusters. Nereus uses a low-overhead controller to select memory-feasible global plans and manages transitions using Elastic Model Units and a global transition graph. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to an initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time, improving end-to-end 8B PPO throughput by 2.14--7.27times over OpenRLHF and 1.10--1.47times over Verl.
Online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling on a real-data trace.
In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time.
Nereus improves end-to-end 8B PPO throughput by 2.14--7.27times over OpenRLHF and by 1.10--1.47times over Verl across diverse clusters.