Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
The authors introduce Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a shared cosine space without updating any text-side backbone parameters. By pairing media samples with dense cascaded captions and using the frozen backbone's own embedding as the teacher target, the student and teacher share identical backbone weights. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner. The 0.9B variant keeps its text weights bit-identical to the backbone, achieving 49.57 nDCG@10 on MTEB-v2 BEIR-8 while being ~2.7x to 9.5x smaller than every open omni embedder compared against.
Omni-Embed-Mini-0.9B achieves 49.57 nDCG@10 on MTEB-v2 BEIR-8.
The model is ~2.7x to 9.5x smaller than every open omni embedder compared against.
The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average.