CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 25 upvotes

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

QUESTION — How can omni reference-to-video generation models be comprehensively evaluated and trained using large-scale structured datasets?

This paper addresses the scarcity of evaluation benchmarks and training data for omni reference-to-video (R2V) generation by introducing OmniVBench and the Omni-R2V Dataset. OmniVBench implements factor-grounded evaluation across 7 task families using 12,172 checklist items to assess reference preservation and disentanglement. Meanwhile, the Omni-R2V Dataset supplies 340K industrial-grade training samples curated from professional video footage via custom pipelines, providing a scalable recipe to support diverse multi-reference R2V data construction for the broader community.

OmniVBench expands R2V evaluation across broader reference types, fine-grained control tasks, and richer reference compositions, covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings.

It introduces factor-grounded evaluation with 12,172 case-specific checklist items, assessing whether intended reference factors are faithfully preserved.

The Omni-R2V Dataset comprises 340K processed training samples spanning diverse reference types and multi-reference compositions.

taesiri · 18 Sept 2026 read the original ↗
↑