CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 48 upvotes

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

QUESTION — How can vision-language models be trained to model physical-world state transitions through interaction?

The paper introduces Spatial-Interactor, a framework that trains vision-language models (VLMs) to model physical-world state transitions through interaction. It organizes learning into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories. The authors construct the LSI-108K dataset and apply a two-stage training strategy using Supervised Fine-Tuning (SFT) followed by On-Policy Distillation (OPD). Experimental results across multiple VLMs and spatial benchmarks demonstrate consistent gains in local transition modeling and long-horizon integration.

Constructed the Learning from Spatial Interaction dataset (LSI-108K) from simulated and real interaction trajectories.

zwq2018 · 19 Sept 2026 read the original ↗
↑