VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
The paper introduces VA-Bench, a benchmark designed to evaluate embodied spatial intelligence through observation, reasoning, action, and revision under incomplete visual evidence. VA-Bench encompasses 14 base task families, geometry/layout variants, and a long-horizon composition track. Evaluating 12 primary model conditions reveals that while top models score high on target localization, their overall macro-average task success remains limited, and active camera control significantly enhances task success compared to passive observation.
The best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run.
The three-run macro-average task success of the best model is only 53.93+/-3.17%.
Active camera control raises task success from 27.86% to 57.50% in a matched comparison.