CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 11 upvotes

Modality-Autoregressive World-Action Models

QUESTION — How can multiple visual modalities be effectively combined within world-action models rather than just predicting RGB?

The paper introduces ModAR, a world-action model (WAM) that autoregressively denoises multiple future modalities before predicting actions. Through training from scratch and evaluating different data mixtures, the authors discover that predicting point tracks, DINO features, and depth maps provides strong benefits, while predicting future RGB alone does not yield consistent gains. ModAR outperforms existing WAM formulations in average success rate. When fine-tuned from the video-model initialization Flex-$\pi$, ModAR achieves a slightly higher observed average success rate while using approximately 20 times fewer training FLOPs and requiring no pretraining.

ModAR achieves the highest average success rate at all evaluated data scales compared to existing WAM formulations.

ModAR achieves a slightly higher observed average success rate (75% vs. 72%) compared to Flex-$\pi$.

ModAR uses approximately 20times fewer training FLOPs and no pretraining when compared against Flex-$\pi$.

taesiri · 15 Sept 2026 read the original ↗
↑