X-Tree: mining reusable skill trees from agent trajectories with zero LLM calls
hllo_wrld · x · 2026-10-06
New work from Waterloo, Duke and NUS: X-Tree. Current agent training (SFT, RLVR) treats trajectories as flat token streams, ignoring the reusable sub-procedures that recur across tasks. Borrowing from text tokenizers, the authors mine a reusable experience tree from existing trajectories with zero LLM calls — scoring action spans by reusability and merging canonicalized actions.
Each node captures how a frequent, success-bearing skill composes from sub-skills, and X-Tree plugs into three training regimes: as data (offline RL training instances), as reward (adaptive skill bonus in online RLVR), and as context (privileged context for a self-teacher in on-policy self-distillation).
Across WebArena, ScienceWorld and WebShop at three model scales, X-Tree improves efficient agent generalization.
More from coding & agent
- Dev Builds MCP Server for memecorp.us So AI Agents Can Post Memes Alongside Humans — AlternativeEast7175 · 2026-10-06
- AI agents are far from mainstream: Muse downloads a fraction of Threads' — FinanceYF5 · 2026-10-06
- Spider-Man 2's traversal physics reverse-engineered into a GTA V mod, mostly built by Opus 5.5 — Promptmethus · 2026-10-06
- Models improve faster than tinkerers: vanilla Codex users get the state of the art — pvncher · 2026-10-06
- Codex keeps hijacking the browser instead of using MCPs, and devs are annoyed — alexgoughcooper · 2026-10-06
- AI-built 3D surf game Tideline goes live: free to play, gamepad-ready, set at Pipeline — majidmanzarpour · 2026-10-06