RASO mines public agent skills as priors, beating TextGrad and GEPA on all four benchmarks

Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation

Jaewon Chu, Ji Soo Lee, Jihwan Park, Dohwan Ko, Jeehye Na, Seunghun Lee, Taehoon Lee, Minseo Yoon, Minseok Joo, Yunyang Xiong, Hyunwoo J. Kim

cs.AI

2026-09-30

RASO retrieves and cross-harness-adapts millions of public agent skills, beating TextGrad and GEPA on all four agent benchmarks at up to 75% lower optimization cost.

What problem this solves

An agent skill is a reusable natural-language document that sits in an agent's context and tells it how to work under a given harness, the execution environment that defines tools, file access, and scoring. Current optimizers (TextGrad, GEPA, SkillOpt, WikiSkill) treat the skill as a text decision variable and rewrite it with rollout feedback on training tasks. That costs twice: rollouts are expensive, and the optimizer only learns from trajectories the agent itself produced, so procedural knowledge beyond its own experience stays out of reach. Meanwhile millions of skills sit publicly shared on GitHub (the GitSkills corpus), spanning diverse tasks and harnesses, and existing methods ignore them.

Plain retrieval is not an answer either. On SpreadsheetBench, 92.9% of retrieved skills come from a different harness than the target, only 7.1% from the right one. Retrieved skills carry source-domain nouns and tool assumptions, so pasting them in injects wrong knowledge: SkillRouter, which reuses retrieved skills as-is, trails even the no-skill baseline on several benchmarks.

Method

RASO wires retrieval from an external skill corpus into both stages of optimization, sharing one pipeline. Skill documents are split into heading-delimited sections and searched with BM25, five sections per query. Cross-Harness Adaptation then rewrites what comes back: an adaptation agent takes the requirement, the task description, the harness description, and the retrieved sections, and emits a short actionable lesson under three rules. Strip source-domain nouns and drop procedures with no counterpart in the target harness. Address only the given requirement. Keep claims about tool or parameter behavior only when the harness description corroborates them. Every lesson ends up phrased in the objects, commands, and units of the target harness.

One worked example: a SpreadsheetBench task needs an empty cell when a department has no qualifying rows. The top-1 retrieved section comes from urban remote sensing, written for numpy raster arrays. Adaptation turns its skip-empty-groups rule into treat-the-output-as-missing, and the generated code leaves the cell blank; without that section, department C's mean lands in B's cell.

RASI (Retrieval-Augmented Skill Initialization) needs zero rollouts: from the task and harness descriptions it generates requirement-query pairs, retrieves, adapts each hit into a lesson, and an initializer agent synthesizes the initial skill. RASU (Retrieval-Augmented Skill Update) works from rollouts: each iteration runs a minibatch with the current skill, derives textual gradients (natural-language critiques of failure modes) plus retrieval queries from failed trajectories, and hands the adapted lessons and gradients to an updater agent. A candidate skill is accepted only if it beats the current one on the validation set.

Results

Initialization at zero rollouts, GPT-5.6-Luna:

MethodOfficeQASpreadsheetALFWorldWebShop
No skill11.4432.9864.4343.56
SkillRouter11.4433.2155.9742.20
RFSI (retrieval-free)40.1144.4069.4043.89
RASI45.7449.1772.6445.06

On Qwen-3.5-9B the RASI-over-RFSI gap widens, up to +10.73 on WebShop (23.27 vs 12.54); SkillRouter manages 9.42 there, below the 16.27 of running with no skill at all.

After full optimization, RASO beats the strongest baseline per benchmark by +3.49 on OfficeQA (49.03 vs 45.54), +6.31 on SpreadsheetBench (63.33 vs 57.02), +1.49 on ALFWorld, and +1.07 on WebShop with GPT-5.6-Luna. On Qwen-3.5-9B the margins grow to +7.47 on ALFWorld and +11.30 on WebShop.

Ablations show the two stages stack: retrieval-free init plus retrieval-free update reaches 40.70/51.67 on OfficeQA/Spreadsheet, RASU alone 47.56/61.07, RASI alone 45.93/58.45, full RASO 49.03/63.33. Cross-Harness Adaptation by itself contributes +3.88 to +7.74 points. K=5 sections per query is the sweet spot, and K=10 degrades. RASO posts the lowest API cost on three of four benchmarks, spending 42-75% less than TextGrad (WebShop: $10.28 vs $41.03) at identical rollout counts.

Why it matters

RASI alone is a deployable cold-start recipe: roughly five points from zero rollouts. The public skill corpus gets wired into the optimization loop for the first time, and even 1% of it already helps while more keeps helping, so the seam is far from exhausted. At matched rollout budgets the total bill lands between a few dollars and a few tens of dollars. The honest caveat: nothing here is exotic, BM25 plus an adaptation agent, and the real engineering weight sits on corpus quality and on the harness description being right.

Limitations

The paper has no dedicated limitations section, so part of this is inferred. Decontamination is keyword-based: the corpus comes from public repos and may contain skills written for the evaluation benchmarks; the authors exclude whole repositories matching benchmark names, which cannot catch benchmark-specific skills that never name the benchmark. The harness description is generated once by a coding agent reading the harness code, and errors there flow straight into the corroboration rule. Training splits are small (39 to 80 tasks) and optimization runs 2 epochs, so whether the component gains survive longer schedules is untested. Some margins sit near noise: +1.49 on ALFWorld and +1.07 on WebShop with GPT-5.6-Luna, and TextGrad and GEPA actually degrade the initial skill on SpreadsheetBench, so shaky baselines flatter the comparison.

Terms

Source

Related papers

All paper explainers