E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
QUESTION — How can few-step generation quality be improved in masked diffusion models without increasing active parameters?
Masked diffusion models generate sequences by progressively unmasking tokens, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime. This study proposes Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts backbone. This approach improves few-step generation over factorized baselines without increasing active parameters.
ArsenyIvanov · 29 Sept 2026
read the original ↗