CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 116 upvotes

Transferring the Intelligence of VLMs to Robotic Control

QUESTION — Can the intelligence of vision-language models generalize from the digital world to the physical world for robotic control?

This paper investigates whether vision-language model intelligence can transfer from digital to physical robotic control via RoboDawn, a human-intuitive interface enabling closed-loop VLM control through compact discrete commands. The authors introduce an in-context learning scheme using a few demonstrations to ground the VLM in interface usage and task strategies. Experiments across RoboTwin 2.0 C2R, RoboDojo, and real-world robots show that the framework achieves strong zero-shot performance and substantial gains in the one-shot setting without requiring task-specific robot training.

On RoboTwin 2.0 C2R, the success rate increases from 53.2% zero-shot to 73.6% one-shot, exceeding the solid baseline π0.5 (46.0%).

Similar gains are observed on RoboDojo, where success rate improves from 35.67% zero-shot to 47.17% one-shot.

The same framework also transfers to real-world robots, performing block-in-basket and block stacking on Franka.

MenghaoGuo · 19 Sept 2026 read the original ↗
↑