CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 20 upvotes

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

QUESTION — How can a text embedding model be extended to multiple modalities without degrading text retrieval quality or requiring billions of parameters?

The authors introduce Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a shared cosine space without updating any text-side backbone parameters. By pairing media samples with dense cascaded captions and using the frozen backbone's own embedding as the teacher target, the student and teacher share identical backbone weights. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner. The 0.9B variant keeps its text weights bit-identical to the backbone, achieving 49.57 nDCG@10 on MTEB-v2 BEIR-8 while being ~2.7x to 9.5x smaller than every open omni embedder compared against.

Omni-Embed-Mini-0.9B achieves 49.57 nDCG@10 on MTEB-v2 BEIR-8.

The model is ~2.7x to 9.5x smaller than every open omni embedder compared against.

The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average.

k-m-irfan · 01 Oct 2026 read the original ↗
↑