CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — system_design 19 upvotes

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training

QUESTION — How can Mixture-of-Experts models be trained at long contexts or large batch sizes without hitting device memory limits from component peak allocations?

Training Mixture-of-Experts (MoE) models at long contexts fails when any component exceeds device memory. The paper identifies four unbounded memory peaks: expert dispatch, vocabulary projection, gradient checkpoint boundaries, and optimizer state. The authors introduce four schedules—PipelinedLLEP, Ring-DTP, Selective Checkpoint Offload (SCO), and OffloadStreamAdamW—to bound these peaks without altering loss or gradients. Tested on 120B to 667B parameter MoE models, these methods cut the MoE dispatch peak by up to 59.3%, the vocabulary projection peak by 86.6%, and speed up the offloaded optimizer step by 2.05times faster, enabling training at 1M context length.

In matched component tests, they cut the MoE dispatch peak by up to 59.3% without losing throughput, the vocabulary projection peak by 86.6%, and the offloaded optimizer step by 2.05times faster.

Composed on MoE models from 120B to 667B parameters, they train at 1M context length, 8--32times the reach of a tuned FSDP2 baseline, and up to 10.4times its throughput.

nxphi47 · 13 Sept 2026 read the original ↗
↑