What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling
The authors investigated why latent WAMs, which discard future generation during inference, fail to retain generalization benefits despite matching explicit WAMs on in-distribution tasks. They found the performance gap arises almost entirely from the first denoising step, meaning the benefit comes from preparing the future rather than generating it. To solve this, they proposed Simple-WAM, which simplifies future modeling into a single forward pass of fully noised video tokens and adapts the noise schedule. Across simulation and real-world tasks, Simple-WAM achieves superior generalization performance while matching the efficiency of latent WAMs.
Latent WAMs fail to retain generalization benefits despite matching explicit ones on in-distribution tasks.
The performance gap arises almost entirely from the first denoising step.
Simple-WAM leads explicit WAMs in generalization performance with efficiency comparable to latent WAMs.