EVOKE elicits pretrained world knowledge in LLM agents via goal-diversity post-training
Yuhan Guo · hf · 2026-10-01
EVOKE is a post-training method for transferable decision-making in LLM agents.
Problem: Agents transfer poorly to unseen environments; world-model approaches that train agents to predict future observations add training cost and compounding planning errors. The authors argue much world knowledge is already internalized from pretraining — the challenge is eliciting it, and typical single-goal post-training provides no pressure to do so, letting policies rely on superficial contextual habits.
Method: Hold environment state and interaction history fixed, then rank the same candidate actions under alternative goals. Theory shows an agent competent across diverse goals must encode a world model recoverable from its action preferences; habit-reliant policies cannot order them correctly, forcing use of pretrained world knowledge.
Results: Across three backbones and diverse tasks, EVOKE improves task performance, unseen-environment generalization, and data efficiency, with controlled analyses explaining the gains.
More from Research
- OSWorld-Science Debuts: 146 Tasks Test How Well VLM Agents Handle Scientific Software — SciAILab · 2026-10-01
- Survey of Attention Evolution: Contextual Memory Becomes the Core of LLM Architecture Design — Zhentao Tan · 2026-10-01
- Hidden Dates in System Prompts Swing LLM Eval Scores by Up to 14% — Mario Sanz-Guerrero · 2026-10-01
- CheatBench Launches to Measure Reward Gaming and Cheating in AI Agents — cais · 2026-10-01
- KLS partially cracked: arXiv paper's core proof ideas generated by AI — burny_tech · 2026-10-01
- Newton's method: when it converges, barely converges, and fails entirely — burny_tech · 2026-10-01