PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation
This work explores vision-and-language navigation (VLN) using panoramic observations and introduces PanoVLN to address why simply replacing perspective images yields limited gains. The authors diagnose that full exploitation requires modifications to action prediction, training supervision, and visual representation. Specifically, they introduce a confidence-guided execution (CGE) strategy for longer-horizon action planning, construct training routes with frequent branching points, and combine semantic and geometric features from RGB panoramas without adding visual tokens. Using a 4B backbone and RGB-only input, PanoVLN outperforms the previous state-of-the-art on R2R-CE and RxR-CE Val-Unseen, and shows faster real-world navigation.
PanoVLN surpasses the previous SOTA by 11.9% and 8.7% in success rate on R2R-CE and RxR-CE Val-Unseen.
The model uses a 4B backbone and RGB-only input.
Real-world experiments on a quadruped demonstrate faster navigation with fewer pauses than prior VLN methods.