X-Tree: mining reusable skill trees from agent trajectories with zero LLM calls

hllo_wrld · x · 2026-10-06

New work from Waterloo, Duke and NUS: X-Tree. Current agent training (SFT, RLVR) treats trajectories as flat token streams, ignoring the reusable sub-procedures that recur across tasks. Borrowing from text tokenizers, the authors mine a reusable experience tree from existing trajectories with zero LLM calls — scoring action spans by reusability and merging canonicalized actions.

Each node captures how a frequent, success-bearing skill composes from sub-skills, and X-Tree plugs into three training regimes: as data (offline RL training instances), as reward (adaptive skill bonus in online RLVR), and as context (privileged context for a self-teacher in on-policy self-distillation).

Across WebArena, ScienceWorld and WebShop at three model scales, X-Tree improves efficient agent generalization.

Original post →

More from coding & agent

coding & agent channel →