Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings
The authors introduce Ovis-Embedding, a state-of-the-art omni-modal embedding family built on a shared multimodal backbone integrating text, image, video, and audio into a common representation space. The approach uses a pretrained Qwen-omni model adapted via contrastive training with low-rank initialization, a curated multi-source data corpus with homogeneous-source sampling, and training/inference optimizations including focal loss and similarity-based Embedding Distillation. At inference time, low-rank feature decomposition yields compact embeddings with flexible dimensionality and minimal performance loss. Empirical evaluations show state-of-the-art performance on MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB.
Ovis-Embedding achieves state-of-the-art performance on MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB.