MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
The paper introduces MotorMind, a robotic manipulation harness that bridges VLM-proposed mid-level actions with deterministic robot control and background memory updates. Operating without task-specific policy training, external coding agents, or grounding tools like SAM3, MotorMind allows a general-purpose vision-language model to perform zero-shot robot manipulation directly from visual observations. Experiments demonstrate significant performance gains over prior zero-shot baselines on LIBERO-PRO suites and real-world xArm6 robot setups under both direct manipulation and human perturbation conditions.
MotorMind achieves 66.7% success on the base LIBERO-PRO suites and 53.8% under perturbations.
The same interface reaches 95% average success on a real xArm6 robot across direct manipulation and human-perturbation settings.