Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes
QUESTION — How can we improve interaction understanding in 3D scenes by coupling geometric and semantic information?
The work presents Segment-Snap, which connects movable parts and handles in 3D scenes through their physical relationship. Learned predictors identify broad part surfaces and small handles, while a geometric decoder uses planar and upright priors to constrain motion. On Articulate3D validation, handle guidance raises motion-gated AP from 13.74% to 40.98% at fixed masks and axes.
Handle guidance raises motion-gated AP from 13.74% to 40.98% at fixed masks and axes.
Additional handle candidates raise handle AP from 24.63% to 29.65%.
Part-based class correction adds 0.98 points.
imsuperkong · 21 Sept 2026
read the original ↗