CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — system_design 9 upvotes

Nereus: Adaptive Parallelism for LLM Post-Training

QUESTION — How can an adaptive runtime adjust execution plans during LLM reinforcement learning post-training to handle changing resource availability and distributed state?

This paper presents Nereus, a cost-aware runtime that adapts LLM reinforcement learning (RL) post-training jobs into efficient execution plans on GPU clusters. Nereus uses a low-overhead controller to select memory-feasible global plans and manages transitions using Elastic Model Units and a global transition graph. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to an initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time, improving end-to-end 8B PPO throughput by 2.14--7.27times over OpenRLHF and 1.10--1.47times over Verl.

Online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling on a real-data trace.

In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time.

Nereus improves end-to-end 8B PPO throughput by 2.14--7.27times over OpenRLHF and by 1.10--1.47times over Verl across diverse clusters.

HollowMan6 · 28 Sept 2026 read the original ↗
↑