CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 73 upvotes

Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings

QUESTION — How can we build a universal omni-modal embedding family supporting text, image, video, and audio in a shared representation space?

The authors introduce Ovis-Embedding, a state-of-the-art omni-modal embedding family built on a shared multimodal backbone integrating text, image, video, and audio into a common representation space. The approach uses a pretrained Qwen-omni model adapted via contrastive training with low-rank initialization, a curated multi-source data corpus with homogeneous-source sampling, and training/inference optimizations including focal loss and similarity-based Embedding Distillation. At inference time, low-rank feature decomposition yields compact embeddings with flexible dimensionality and minimal performance loss. Empirical evaluations show state-of-the-art performance on MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB.

Ovis-Embedding achieves state-of-the-art performance on MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB.

taesiri · 21 Sept 2026 read the original ↗
↑