SkillLift cuts agent skill-evolution token cost 40-70% by ranking, not rollouts

dair_ai · x · 2026-09-22

A paper introduced by dair-ai shows agents can improve their skill prompts by ranking candidates with a learned rubric instead of scoring every revision with a full rollout — the rollout cost is why skill self-evolution usually only patches observed failures. SkillLift trains a rubric to agree with real outcomes on pairwise skill rankings; an inner loop revises skills against the frozen rubric at no rollout cost, while an outer loop spends a few real rollouts to re-align the rubric by rank correlation. On SkillsBench and WildClawBench (147 tasks) across three models, it beats SkillOpt and CoEvoSkills in all six combinations, with 40–70% less token cost than frontier evolving methods.

Original post →

More from coding & agent

coding & agent channel →