Grounded Action Model: 3D Grounding as a Foundation for Robotics
This paper proposes Grounded Action Models (GAM), a robot foundation model paradigm built around explicit 3D grounding. GAM conditions on language, point, or box prompts, transforming them into an object-centric representation combining visual features and metric geometry, which is then mixed with robot state history via a multi-stream transformer to predict action chunks. Evaluated on RoboTwin 2.0, GAM achieves an average success rate of 55.3% across 50 tasks versus 52.0% for Spatial Forcing, and achieves a state-of-the-art average success rate of 61% on LIBERO-PRO across 16 perturbation settings, while retaining robust performance on real hardware under visual shifts.
GAM achieves an average success rate of 55.3% across 50 tasks (vs. 52.0% for Spatial Forcing).
47.6% under scene randomization (vs. 30.4% for Abot-M0).
achieves a state-of-the-art average success rate of 61% (vs. 53% for π_{0.5}) across 16 perturbation settings.