WorldLine: Action-Driven Visual Simulation for Robotic Manipulation
The research presents WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine trains manipulation dynamics on over 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments, utilizing an image-space action representation as a shared interface. Robot-focused few-step distillation enables efficient causal rollouts while retaining motion accuracy. On failed trajectories, WorldLine improves robot-mask IoU by 0.1626 over the strongest baseline, predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, and improves task success by up to 21.4 percentage points over direct policy execution.
WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments.
On failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline.
It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, improving task success by up to 21.4 percentage points over direct policy execution.