IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
The paper proposes IntBMoE, a block-conditioned Mixture-of-Experts architecture that decouples token participation, compute execution, and parameter materialization. By pairing dense expert composition with sparse block execution via a learned codebook and a lightweight hypernetwork, IntBMoE achieves full participation while keeping execution sparse and memory bounded. Dual-Path Residual Gating (DPRG) further couples paths via multiplicative gating. IntBMoE is fully deployed in AMap's generative recommendation system, serving hundreds of millions of users under a 60ms latency budget with a 2.4% relative UVCTR gain.
IntBMoE decouples token participation, compute execution, and parameter materialization in MoE architectures.
The model is fully deployed in AMap's generative recommendation system under a 60ms latency budget.
Online A/B testing demonstrates a 2.4% relative UVCTR gain.