EVISKILL: replayable evidence cards ground continual skill evolution in LLM agents

Yan Zhou · hf · 2026-10-07

The paper introduces EVISKILL, addressing lost behavioral evidence when LLM agents accumulate reusable procedural knowledge from interaction. Prior experience-driven methods can drop the evidence and task contexts supporting edits, and global validation poorly judges local changes. EVISKILL organizes execution observations into Replayable Evidence Cards with explicit links from edits to supporting contexts; targeted replay re-executes to verify edits and feed corrections. Across epochs, supported edits are provisionally retained for refinement, while global validation governs incorporation into the final skill.

Evaluated on three interactive benchmarks across six LLM backbones, showing consistent effectiveness.

Original post →

More from coding & agent

coding & agent channel →