CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 5 upvotes

RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs

QUESTION — How can open-vocabulary relation prediction be performed in real time from arbitrary inputs without being constrained by fixed object labels?

The authors introduce RelateAnything, a 53M-parameter model that takes an image and regions from any source to return scored relations over a predicate vocabulary supplied as strings at inference. The model utilizes a bank of text embeddings instead of a learned classifier and runs at 20 ms/frame. To enable training, the authors constructed RA-4M with 474k images and 4.3M relations, and established OV-SGG-Bench to evaluate performance across diverse axes.

RelateAnything is a 53M-parameter model running at 20 ms/frame.

RA-4M comprises 474k images and 4.3M relations over 10,102 free-text predicates.

On three cross-dataset benchmarks and a fourth zero-shot, RelateAnything has 2.3-3.5x the mean recall of the strongest open-vocabulary method of comparable scale.

nielsr · 11 Sept 2026 read the original ↗
↑