Learning Foresight without Explicit Trajectories for 3D Diffusion Policies
The authors introduce Movement Trend Guidance, a method allowing 3D diffusion policies to capture interaction foresight without explicit trajectory planning. From a short observation history, the policy learns a compact latent representation of interaction evolution supervised by sparse future gripper states. This latent provides global conditioning for action generation alongside a gated FiLM branch at the UNet bottleneck. Adding only 3.52% more parameters to DP3, the method consistently improves performance across RoboTwin2.0, LIBERO-40, and real-robot tasks.
The method adds only 3.52% more parameters to DP3.
It reaches 62.8% vs. 56.1% in 50-task RoboTwin2.0 mixed training.
It achieves 71.93% vs. 37.08% on LIBERO-40 and 72.0% vs. 49.0% on five real-robot tasks.