SkillsBench accepted at NeurIPS 2026: 87-task benchmark tests whether agent skills actually help

HanchungLee · x · 2026-09-25

SkillsBench has been accepted to the NeurIPS 2026 Evaluations and Datasets Track as a poster. The benchmark asks whether agent skills actually improve LLM agent performance: it spans 87 tasks across 8 domains, each paired with curated skill packages and deterministic verifiers, and runs matched no-Skills vs curated-Skills comparisons to quantify gains. Co-authored by Dawn Song, Han-chung Lee, and 76 others.

Original post →

More from coding & agent

coding & agent channel →