CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — architecture 60 upvotes

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

QUESTION — How can few-step generation quality be improved in masked diffusion models without increasing active parameters?

Masked diffusion models generate sequences by progressively unmasking tokens, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime. This study proposes Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts backbone. This approach improves few-step generation over factorized baselines without increasing active parameters.

ArsenyIvanov · 29 Sept 2026 read the original ↗
↑