X-Tree tokenizes reusable experience into skill hierarchies, boosting agent RL by up to 5.8%

UWaterloo · hf · 2026-10-03

UWaterloo's X-Tree recovers reusable skill hierarchies from agent trajectories without LLM calls: action spans are scored by reusability and merged into a canonicalized experience tree, following the vocabulary-building logic of text tokenizers. It integrates into offline RL (nodes as training instances), online RLVR (adaptive skill bonus), and on-policy self-distillation (privileged teacher context). At matched data and budget across three model scales, it improves up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld, and 4.1% success on WebShop, with ablations attributing gains to the tree structure.

Original post →

More from coding & agent

coding & agent channel →