MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
This paper introduces Mixture of Memory Embeddings (MoME), a conditional context-aware memory mechanism that replaces a static token embedding table with a mixture of M slots. Using a learned gate over hidden states to route at each position, MoME resolves polysemous tokens into distinct memory slots depending on context. Controlled pretraining experiments across multiple backbone architectures demonstrate that MoME improves over Value Embedding, Bigram, and STEM baselines under iso-parameter and iso-training-FLOP settings while showing promising scaling trends at sub-billion scale.
It improves over Value Embedding, Bigram, and STEM baselines in iso-parameter and iso-training-FLOP settings.
It shows a more promising memory-size scaling trend at sub-billion scale.