CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 20 upvotes

HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

QUESTION — How can perception, retrieval, and navigation be coordinated in embodied navigation tasks without retraining the underlying model?

The study introduces HarnessVLN, a zero-shot, training-free framework that uses an Agent Harness to coordinate perception, retrieval, grounding, navigation, recovery, and termination through a unified tool interface. The system validates planner proposals against spatial evidence and geometric feasibility, while maintaining a Spatiotemporal Graph for reusable spatial evidence. Results show that HarnessVLN achieves success rates of 60.8%, 53.9%, 76.0%, and 59.3% on R2R, RxR, HM3D-v2, and HM3D-OVON, respectively, surpassing prior training-free SOTA results.

HarnessVLN is a zero-shot, training-free framework integrating an Agent Harness that coordinates through a unified tool interface.

A Spatiotemporal Graph maintains reusable spatial evidence and failure annotations for verification and recovery.

HarnessVLN achieves success rates of 60.8%, 53.9%, 76.0%, and 59.3% on R2R, RxR, HM3D-v2, and HM3D-OVON, respectively.

chenyang0126 · 14 Sept 2026 read the original ↗
↑