CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 37 upvotes

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

QUESTION — How effectively do general-purpose multimodal large language models execute the complete observe-reason-act-revise loop for embodied spatial intelligence?

The paper introduces VA-Bench, a benchmark designed to evaluate embodied spatial intelligence through observation, reasoning, action, and revision under incomplete visual evidence. VA-Bench encompasses 14 base task families, geometry/layout variants, and a long-horizon composition track. Evaluating 12 primary model conditions reveals that while top models score high on target localization, their overall macro-average task success remains limited, and active camera control significantly enhances task success compared to passive observation.

The best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run.

The three-run macro-average task success of the best model is only 53.93+/-3.17%.

Active camera control raises task success from 27.86% to 57.50% in a matched comparison.

MrBean2024 · 17 Sept 2026 read the original ↗
↑