Scaling Laws for Looped Mixture of Experts
They introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping characterizing the effective-parameter gain from looping. The laws predict held-out loss more accurately than prior alternatives and recover standard dense and MoE scaling laws as special cases. Downstream evaluations demonstrate that sparsity delivers ~3x active-parameter efficiency, while recurrence yields ~2x total-parameter efficiency on reasoning. At trillion-token scale, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on reasoning benchmarks while enabling test-time scaling.
Sparsity delivers ~3x active-parameter efficiency.
Recurrence yields ~2x total-parameter efficiency on reasoning.
At trillion-token scale, a looped MoE matches a ~2x larger non-looped MoE on reasoning benchmarks.