Meta paper: Agents must learn when to use tools, not just have them, achieving 96.9% on ALFWorld
daniel_mac8 · x · 2026-08-24
A new Meta paper suggests that simply giving an agent more memory and tools is insufficient; the agent must learn when they are worth using. The proposed EvoHarness-RL method achieved a 96.9% success rate on ALFWorld using Qwen2.5-8B, while reducing harness usage to about 1 call per episode, addressing state management and forgetfulness in long-running tasks.
Related event: Meta's EvoHarness-RL Trains Agents to Decide When to Use Memory(2 posts)→
More from coding & agent
- DAIR launches free hands-on lab for Exo agent framework — omarsar0 · 2026-08-24
- Open-source Exo framework enables agent self-modification and time travel rollback — omarsar0 · 2026-08-24
- What should human approval bind to when agent resubmission changes the request hash? — docybo · 2026-08-24
- Enterprise Agents Need Architecture Constraints, Not Just Data Quality — jonerp · 2026-08-24
- Google launches Developer Knowledge MCP integrated into gcloud CLI — rseroter · 2026-08-24
- Dual RTX 6000s run Qwen3.8-27B at 150t/s, yet 12x slower than Claude on the task — EkbatDeSabat · 2026-08-24