Beyond the Current Scene: Event-Referential Grasping with Active View Selection
They introduce BeyondSCe, a zero-shot robotic system that handles requests specifying targets by their role in a past event rather than by name or appearance. When targets are occluded, it combines an event prior recovered from history with current scene geometry to select camera viewpoints likely to reveal the target. Using pretrained models without task-specific training, it achieves grasp success rates of 76% for initially visible and 77% for occluded targets, and increases success rates from 75% to 95% on heavily occluded scenes while reducing the mean number of views from 3.35 to 2.20.
Achieves grasp success rates of 76% and 77% for initially visible and occluded targets, respectively.
Increases grasp success rates from 75% to 95% on four additional scenes with heavy occlusion.
Reduces the mean number of views from 3.35 to 2.20 compared to an active-perception baseline.