World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
This work presents World Action Agent (WAA), a multi-agent harness enabling vision-language models (VLMs) to pilot robots inside a visual action workspace featuring contact views, action rehearsal, and in-view correction. WAA evolves multimodal skills from expert demonstrations, consults them via a Skill Agent, and uses interaction traces to fine-tune smaller VLMs. Evaluated on LIBERO-Pro, WAA achieves a state-of-the-art average success rate using skills evolved exclusively from LIBERO-90, and maintains effectiveness on robosuite without further learning. Furthermore, fine-tuning Qwen3.5-9B on harness traces increases its out-of-domain success rate significantly.
On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.
Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.