Transferring the Intelligence of VLMs to Robotic Control
This paper investigates whether vision-language model intelligence can transfer from digital to physical robotic control via RoboDawn, a human-intuitive interface enabling closed-loop VLM control through compact discrete commands. The authors introduce an in-context learning scheme using a few demonstrations to ground the VLM in interface usage and task strategies. Experiments across RoboTwin 2.0 C2R, RoboDojo, and real-world robots show that the framework achieves strong zero-shot performance and substantial gains in the one-shot setting without requiring task-specific robot training.
On RoboTwin 2.0 C2R, the success rate increases from 53.2% zero-shot to 73.6% one-shot, exceeding the solid baseline π0.5 (46.0%).
Similar gains are observed on RoboDojo, where success rate improves from 35.67% zero-shot to 47.17% one-shot.
The same framework also transfers to real-world robots, performing block-in-basket and block stacking on Franka.