World Embedding Benchmark
The authors introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases across 80 families, to evaluate how video representations encode physical information. They reveal a trade-off between cross-modal physical alignment and quantitative information recoverability. Utilizing these embeddings to retrieve reference videos for retrieval-augmented generation improves the physical fidelity of generated videos with models like MiniMax-H3.
The benchmark comprises 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics & electromagnetism.
Continual contrastive training improves retrieval and pair classification but degrades physical-property regression.
Retrieved references improve the physical fidelity of generated videos with MiniMax-H3.