X-Tree tokenizes reusable experience into skill hierarchies, boosting agent RL by up to 5.8%
UWaterloo · hf · 2026-10-03
UWaterloo's X-Tree recovers reusable skill hierarchies from agent trajectories without LLM calls: action spans are scored by reusability and merged into a canonicalized experience tree, following the vocabulary-building logic of text tokenizers. It integrates into offline RL (nodes as training instances), online RLVR (adaptive skill bonus), and on-policy self-distillation (privileged teacher context). At matched data and budget across three model scales, it improves up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld, and 4.1% success on WebShop, with ablations attributing gains to the tree structure.
More from coding & agent
- Dev turns a Pixel 10 Pro XL into an offline OpenAI-compatible server running Gemma 4 E4B at ~11 tok/s — VerityAISolutions · 2026-10-03
- NVIDIA hands over first Vera CPU to test AI agent code environments — denisyarats · 2026-10-03
- TL: a token-efficient language that cuts LLM code tokens by 30-60% — Ok-Condition7148 · 2026-10-03
- Context compaction backfires: Billion Context plugin sends coding agent into a loop — Iory1998 · 2026-10-03
- A weekend, a few hundred lines: building a personal AI voice agent with Telnyx, gpt-live and Cloudflare — itsOmSarraf_ · 2026-10-03
- Mnemos plugin to visualize agent memory in your platform's canvas — RileyRalmuto · 2026-10-03