RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents
This paper demonstrates that the impressive success rates of modern embodied agents often stem from low scene entropy, where visual cues alone suffice and language instructions are ignored. The authors introduce RoboFollow, a diagnostic benchmark featuring high scene entropy, a four-level hierarchical diagnostic protocol (L0-L3), and confound-controlled diagnosis. Evaluating nine VLA and WAM policies reveals that strong L0 baseline performance does not reliably transfer to L1-L3, and standard mitigation strategies fail to bridge this instruction-following gap.
Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup.
Representative mitigations, including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance, all fail to close this gap.