UniSkill: Actor-Aligned Skill Proposals Hit 98.4% Success on ALFWorld
Yifei Lu · hf · 2026-10-08
A new paper, UniSkill, introduces a method for LLM agents to distill reusable skills from past interactions.
- Problem: Prior work jointly optimizes task execution and skill extraction, but rewards skill proposals via reuse can conflate skill benefits with actor improvement, while testing each proposal directly requires costly extra rollouts.
- Method: A shared policy interacts with the environment and proposes skillbank edits (Add/Update/No Edit) from resulting trajectories. The actor learns from environment rewards, while contrastive action feedback — measuring how swapping in a proposed skill changes the actor's action log-likelihood gap between prior successful and failed trajectories — provides an actor-alignment signal without new rollouts. Skill-edit support regularization preserves exploration.
- Results: 98.4% success on ALFWorld and 84.7% on WebShop with stable joint training; remains effective with smaller backbones.
- Implementation is open-sourced at GitHub (LimOkii/UniSKill).
More from coding & agent
- Harvard study of 718 firms finds AI coding boosts individual output ~30% but firm-level shipping doesn't move — DavidLinthicum · 2026-10-08
- WSJ columnist declares 'the year of Linux' is finally here, thanks to AI agents — alexvoica · 2026-10-08
- Full materials from a 2-day workshop on using Claude Code and Codex for academic research — JeremyNguyenPhD · 2026-10-08
- Attestation: local-first MCP tools turn research provenance into CI-enforceable claims — JeremyCMorgan · 2026-10-08
- Tensorlake npm SDK 0.5.144 Found Shipping Credential-Stealing Malware — TechNadu · 2026-10-08
- Enterprise agent lesson: permissions don't guarantee correct business transactions — shashib · 2026-10-08