HappyWorld-Bench
The paper introduces HappyWorld-Bench, a comprehensive benchmark evaluating world models across a hierarchical framework of six capabilities (W1-W6) across three independent tracks: video, spatial, and embodied world models. It comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Operating HappyWorld-Arena for human A/B comparisons yields Elo ratings that complement automated behavioral metrics. Evaluating multiple models reveals reliability gaps: video models show reduced consistency during extended rollouts, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions.
Spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution.
HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases.