RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs
The authors introduce RelateAnything, a 53M-parameter model that takes an image and regions from any source to return scored relations over a predicate vocabulary supplied as strings at inference. The model utilizes a bank of text embeddings instead of a learned classifier and runs at 20 ms/frame. To enable training, the authors constructed RA-4M with 474k images and 4.3M relations, and established OV-SGG-Bench to evaluate performance across diverse axes.
RelateAnything is a 53M-parameter model running at 20 ms/frame.
RA-4M comprises 474k images and 4.3M relations over 10,102 free-text predicates.
On three cross-dataset benchmarks and a fourth zero-shot, RelateAnything has 2.3-3.5x the mean recall of the strongest open-vocabulary method of comparable scale.