CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 8 upvotes

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

QUESTION — How can visual goal and dynamics prediction be integrated into robot action generation through a shared trajectory model?

This paper presents Dynin-Robotics, an omnimodal masked-diffusion model that integrates visual goal and dynamics prediction into robot action generation via a shared trajectory model. By varying conditioning and target spans, the model handles action prediction, action-conditioned next-observation prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction. Continually pretrained on approximately 1.33 million trajectories from 48 Open X-Embodiment datasets, the system achieves a 78.4% average success rate on a Franka Research 3 robot and accelerates model-side action decoding by up to 29.2x through an optimized block-parallel implementation.

It is continually pretrained on approximately 1.33 million trajectories from 48 Open X-Embodiment datasets.

It achieves a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot.

An optimized block-parallel implementation accelerates model-side action decoding by up to 29.2x relative to the base implementation.

leehe228 · 11 Sept 2026 read the original ↗
↑