In-Context Robot Learning with VLM Agents
The paper introduces GPT-Policy, a general-agent framework for in-context robot learning using a vision-language model. The system integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action. Real-robot trials show that human video demonstrations improve task completion even without robot action labels, establishing an empirical foundation for translating the general capabilities of VLMs into physical behavior.
GPT-Policy is a framework integrating a context compiler, VLM, and constrained controller for in-context robot learning.
Human video demonstrations improve task completion even without robot action labels in real-robot trials.
Aligned action references yield further gains on contact-sensitive tasks.