Modality-Autoregressive World-Action Models
The paper introduces ModAR, a world-action model (WAM) that autoregressively denoises multiple future modalities before predicting actions. Through training from scratch and evaluating different data mixtures, the authors discover that predicting point tracks, DINO features, and depth maps provides strong benefits, while predicting future RGB alone does not yield consistent gains. ModAR outperforms existing WAM formulations in average success rate. When fine-tuned from the video-model initialization Flex-$\pi$, ModAR achieves a slightly higher observed average success rate while using approximately 20 times fewer training FLOPs and requiring no pretraining.
ModAR achieves the highest average success rate at all evaluated data scales compared to existing WAM formulations.
ModAR achieves a slightly higher observed average success rate (75% vs. 72%) compared to Flex-$\pi$.
ModAR uses approximately 20times fewer training FLOPs and no pretraining when compared against Flex-$\pi$.