CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — architecture 17 upvotes

Scaling Laws for Looped Mixture of Experts

QUESTION — How can recurrence and sparsity be jointly modeled to scale looped Mixture-of-Experts models efficiently?

They introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping characterizing the effective-parameter gain from looping. The laws predict held-out loss more accurately than prior alternatives and recover standard dense and MoE scaling laws as special cases. Downstream evaluations demonstrate that sparsity delivers ~3x active-parameter efficiency, while recurrence yields ~2x total-parameter efficiency on reasoning. At trillion-token scale, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on reasoning benchmarks while enabling test-time scaling.

Sparsity delivers ~3x active-parameter efficiency.

Recurrence yields ~2x total-parameter efficiency on reasoning.

At trillion-token scale, a looped MoE matches a ~2x larger non-looped MoE on reasoning benchmarks.

yanbeic · 30 Sept 2026 read the original ↗
↑