EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making
This paper argues that typical post-training fails to elicit internalized world knowledge in LLMs, leading to poor generalization in unseen environments. To overcome this without training explicit world models, the authors introduce EVOKE, a post-training method that introduces goal diversity at fixed states. By ranking candidate actions under alternative goals while keeping the environment state constant, EVOKE prevents policies from relying on superficial contextual habits. This forces the model to implicitly elicit its pretrained world knowledge for correct decision-making. Evaluations across diverse tasks and backbone models demonstrate improved performance, better zero-shot generalization to unseen environments, and enhanced data efficiency.
EVOKE demonstrates improved task performance and unseen environment generalization.
The method achieves enhanced data efficiency through direct decision supervision under goal diversity.