SkillsBench accepted at NeurIPS 2026: 87-task benchmark tests whether agent skills actually help
HanchungLee · x · 2026-09-25
SkillsBench has been accepted to the NeurIPS 2026 Evaluations and Datasets Track as a poster. The benchmark asks whether agent skills actually improve LLM agent performance: it spans 87 tasks across 8 domains, each paired with curated skill packages and deterministic verifiers, and runs matched no-Skills vs curated-Skills comparisons to quantify gains. Co-authored by Dawn Song, Han-chung Lee, and 76 others.
More from coding & agent
- TypeLLM adds type-safe, schema-guaranteed generation to LLMs without touching weights — kalyan_kpl · 2026-09-25
- Five mistakes people make with agentic factories: micromanaging, no quality bar, and 'puppet shows' — hugobowne · 2026-09-25
- Open Claude Code skills catalog ships 31 testing skills, from metamorphic testing to E2E — doodlestein · 2026-09-25
- Researcher proposes agent-suggests-human-executes loop for real-world science experiments — suragnair · 2026-09-25
- CrawlRaven hits 700 users in 60 days, targets 2,000 more with MCP-powered SEO audits — ayushtweetshere · 2026-09-25
- Human-in-the-loop or machine-executed: verifiable tasks turn agent traces into RL rewards — suragnair · 2026-09-25