CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 18 upvotes

InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

QUESTION — How can large-scale pretrained priors of visual dynamics, scene semantics, and geometry be unified into a generalist robot action model?

The authors introduce InternW0-Δ, a World Action Model (WAM) that combines visual dynamics, semantics, 4D geometry, and action generation via a Mixture-of-Transformers (MoT) architecture. It uses a pretrained video expert and an action expert interacting under frozen VLM semantic guidance, along with Causal Imprint to supply future scene changes directly to the action expert without future-video rollout at inference. The model is pretrained on a heterogeneous corpus containing over 20K hours of processed training data.

The resulting corpus contains over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind.

taesiri · 25 Sept 2026 read the original ↗
↑