Peking University Introduces ContinualSkillBench: Evaluating Continual Skill Evolution in LLM Agents

PekingUniversity · hf · 2026-08-05

Peking University introduced ContinualSkillBench, a dynamic evaluation framework designed to test whether LLM agents can genuinely evolve their capabilities using external skill libraries.

The benchmark covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Key findings from their experiments include:

Related event: New Benchmark Shows Explicit Skill Libraries Fail to Significantly Boost AI Agents(3 posts)→

Original post →

More from coding & agent

coding & agent channel →