CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 45 upvotes

HappyWorld-Bench

QUESTION — How can we comprehensively evaluate the consistency and responsiveness of world models as agents interact and explore?

The paper introduces HappyWorld-Bench, a comprehensive benchmark evaluating world models across a hierarchical framework of six capabilities (W1-W6) across three independent tracks: video, spatial, and embodied world models. It comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Operating HappyWorld-Arena for human A/B comparisons yields Elo ratings that complement automated behavioral metrics. Evaluating multiple models reveals reliability gaps: video models show reduced consistency during extended rollouts, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions.

Spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution.

HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases.

CheeryLJH · 21 Sept 2026 read the original ↗
↑