HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness
The study introduces HarnessVLN, a zero-shot, training-free framework that uses an Agent Harness to coordinate perception, retrieval, grounding, navigation, recovery, and termination through a unified tool interface. The system validates planner proposals against spatial evidence and geometric feasibility, while maintaining a Spatiotemporal Graph for reusable spatial evidence. Results show that HarnessVLN achieves success rates of 60.8%, 53.9%, 76.0%, and 59.3% on R2R, RxR, HM3D-v2, and HM3D-OVON, respectively, surpassing prior training-free SOTA results.
HarnessVLN is a zero-shot, training-free framework integrating an Agent Harness that coordinates through a unified tool interface.
A Spatiotemporal Graph maintains reusable spatial evidence and failure annotations for verification and recovery.
HarnessVLN achieves success rates of 60.8%, 53.9%, 76.0%, and 59.3% on R2R, RxR, HM3D-v2, and HM3D-OVON, respectively.