Meta paper: Agents must learn when to use tools, not just have them, achieving 96.9% on ALFWorld

daniel_mac8 · x · 2026-08-24

A new Meta paper suggests that simply giving an agent more memory and tools is insufficient; the agent must learn when they are worth using. The proposed EvoHarness-RL method achieved a 96.9% success rate on ALFWorld using Qwen2.5-8B, while reducing harness usage to about 1 call per episode, addressing state management and forgetfulness in long-running tasks.

Related event: Meta's EvoHarness-RL Trains Agents to Decide When to Use Memory(2 posts)→

Original post →

More from coding & agent

coding & agent channel →