CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 7 upvotes

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

QUESTION — How can vision-language models (VLMs) be integrated into a visual action workspace to pilot robot manipulation more effectively?

This work presents World Action Agent (WAA), a multi-agent harness enabling vision-language models (VLMs) to pilot robots inside a visual action workspace featuring contact views, action rehearsal, and in-view correction. WAA evolves multimodal skills from expert demonstrations, consults them via a Skill Agent, and uses interaction traces to fine-tune smaller VLMs. Evaluated on LIBERO-Pro, WAA achieves a state-of-the-art average success rate using skills evolved exclusively from LIBERO-90, and maintains effectiveness on robosuite without further learning. Furthermore, fine-tuning Qwen3.5-9B on harness traces increases its out-of-domain success rate significantly.

On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.

Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.

taesiri · 24 Sept 2026 read the original ↗
↑