Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
The paper introduces Spatial-Interactor, a framework that trains vision-language models (VLMs) to model physical-world state transitions through interaction. It organizes learning into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories. The authors construct the LSI-108K dataset and apply a two-stage training strategy using Supervised Fine-Tuning (SFT) followed by On-Policy Distillation (OPD). Experimental results across multiple VLMs and spatial benchmarks demonstrate consistent gains in local transition modeling and long-horizon integration.
Constructed the Learning from Spatial Interaction dataset (LSI-108K) from simulated and real interaction trajectories.