ExpVoyager: On-Demand Trajectory Navigation Lifts ALFWorld Success from 41% to 64%

ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis

Kwangwook Seo, Dongha Lee

cs.CL

2026-09-26

ExpVoyager navigates raw agent trajectories on demand to write a skill.md. On Qwen3.5-9B, ALFWorld success is 63.7% versus 51.1% for SkillTTA and 41.4% for ReAct.

What problem this solves

Agent skills used to be hand-written playbooks. The newer move is to distill reusable procedures from the agent's own traces. AWM, ReasoningBank, and Trace2Skill all freeze that distillation before anyone knows which downstream task will arrive: they store a workflow or a memory entry, then retrieve it later.

The freeze is early, and it throws knowledge away. Yonsei built 80 oracle skills on ALFWorld whose helpfulness was checked by running a frozen executor, then labeled supporting knowledge in 300 source traces. Pre-built artifacts keep at most 40.9% of that oracle knowledge (AWM). Trace2Skill keeps 31.5%. ReasoningBank keeps 12.7%. Retrieval does not recover the rest. Raising k lifts knowledge recall a little and drops precision fast. Similar traces carry irrelevant local detail. A cue that matters can sit in a trace that does not look similar.

The question is whether the agent can walk raw traces for the current task, at the resolution that task needs.

Method

ExpVoyager treats skill synthesis as experience navigation. The executor θ stays frozen. A skill curator ϕ searches the experience pool E and writes a skill.md for the current task x, then passes it to the executor as extra context.

The curator does not retrieve a whole trace as one blob. It uses a Navigable Interface that exposes two views of the same record:

Each round picks one action. searchexp takes a view level, fields, a regex, and a hit cap, and returns complete step or trajectory records plus a source pointer. inspecttraj expands that pointer into the chronological trace. The default budget is 20 rounds, with an early FINISH.

Navigation State Z=(K,Q) is updated after every observation. Each item in K splits three ways: what happened in the source, which procedural relation might transfer, and what still has to be checked in the current environment. Q holds open questions, seeded from the task and rewritten as evidence arrives. The state works like a question list carried through lab notebooks: each new record revises confirmed procedures and remaining gaps. After the last round, the curator compiles Z into skill.md: transferable steps, runtime checks for unresolved conditions, and recovery notes from failures.

Similarity search treats surface likeness as usefulness. The interface lets the model choose which layer to read and which trace to follow. The state stops it from treating a newspaper on armchair 2 in some old episode as a fact about the current room.

Results

Default run: Qwen3.5-9B as both curator and executor, 1000 offline traces on ALFWorld and WebShop, 500 on ScienceWorld, mean of three random seeds.

MethodALFWorld SRWebShop SRScienceWorld SR
ReAct41.419.627.0
AWM45.624.232.6
ReasoningBank43.520.329.6
Trace2Skill45.422.631.9
SkillTTA51.119.229.6
ExpVoyager63.728.737.4

ALFWorld steps fall from 34.1 (ReAct) to 25.9. WebShop score is 61.2 against SkillTTA's 34.3 and ReAct's 50.3. ScienceWorld score is 47.0 against AWM's 42.7, the strongest pre-built baseline there.

With a Gemma4-31B executor and the same 9B curator, ALFWorld success moves from 53.9 to 69.6. A 31B curator reaches 71.3. GPT-5.4-mini as curator reaches 75.9. A small curator already lifts a much larger executor, so picking what to reuse and solving the task are separate jobs.

In the online setting there is no pre-collected pool, only traces from earlier tasks in the stream. Gains are small or negative when experience is scarce. By the end of the stream, cumulative success gain over ReAct is about 15% on ALFWorld (baselines stall around 3% to 7%) and about 18% on ScienceWorld (baselines around 6% to 11%). Offline, nested subsets up to 1000 traces make baselines worse and ExpVoyager better.

ALFWorld ablation: drop the Navigable Interface and success falls 63.7 to 55.3, with 52% of found knowledge marked irrelevant. Drop Navigation State and success is 58.9, with 47% duplicated. The full system marks about 62% of finds as new. Extra rounds help up to 20 and then hurt. Under matched token budgets, multi-round agentic search on AWM, ReasoningBank, and SkillTTA still lags. Starting from a pre-built skill and navigating to patch it usually raises score and cuts end-to-end tokens, with one exception reported. End-to-end cost is about 270k to 450k tokens per task.

Why it matters

This is a fork in the skill layer of an agent harness. skill.md does not have to be written before deployment; it can be compiled from raw traces at task time. A stack that already has AWM or ReasoningBank can keep those files as a starting point and let ExpVoyager fill the hole for the current task.

The bill is real. Twenty navigation rounds and 270k to 450k tokens per task is an order of magnitude above retrieving three workflows. The setup fits interactive environments with clean trace structure. There is no experiment on messy tool-call logs.

The result is about access. More experience makes pre-distilled skills noisier and navigation stronger.

Limitations

There is no Limitations section. The bounds sit in the protocol and in the failure modes the authors already report.

All three benchmarks are text games: household, shopping, science lab, capped at 50, 15, and 30 steps. No software engineering, live websites, or long tool chains. Oracle labeling and claim matching used a stronger model plus human review, but only on 80 ALFWorld oracle skills and 300 source traces. The claim that more than half the knowledge is gone does not automatically transfer to WebShop.

Sparse experience can hurt. Past 20 rounds, extra navigation also hurts. The loop is sensitive to noise and over-search. The executor is frozen and the skill is context-only; environment reward never trains the curator. On WebShop, SkillTTA's 19.2 SR and 34.3 score both sit below ReAct, so part of the gap is a weak test-time baseline. Failure cases appear in the appendix without a typed breakdown. There is no equal-accuracy comparison against simply sampling ReAct more times.

Terms

Source

Related papers

All paper explainers