Peking University Introduces ContinualSkillBench: Evaluating Continual Skill Evolution in LLM Agents

PekingUniversity · hf · 2026-08-05

Peking University introduced ContinualSkillBench, a dynamic evaluation framework designed to test whether LLM agents can genuinely evolve their capabilities using external skill libraries.

The benchmark covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Key findings from their experiments include:

Original post →

More from coding & agent

coding & agent channel →