CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 13 upvotes

RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

QUESTION — Do modern embodied agents actually follow language instructions, or do they exploit low scene entropy?

This paper demonstrates that the impressive success rates of modern embodied agents often stem from low scene entropy, where visual cues alone suffice and language instructions are ignored. The authors introduce RoboFollow, a diagnostic benchmark featuring high scene entropy, a four-level hierarchical diagnostic protocol (L0-L3), and confound-controlled diagnosis. Evaluating nine VLA and WAM policies reveals that strong L0 baseline performance does not reliably transfer to L1-L3, and standard mitigation strategies fail to bridge this instruction-following gap.

Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup.

Representative mitigations, including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance, all fail to close this gap.

Mqleet · 22 Sept 2026 read the original ↗
↑