SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning
The authors introduce SpatialSpeak, a two-stage framework connecting QA-native reconstruction pretraining with spatial chain-of-thought learning for vision-language models. The first stage combines fine-grained local geometry queries with global scene context queries, while the second stage trains the model to derive answers using spatial chain-of-thought with visual compensation. On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and SpatialSpeak achieves a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.
On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points.
SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.