SkillGym builds 6.8K verifiable skill-use environments; 9B SFT model beats 397B
Renxi Wang · hf · 2026-10-01
SkillGym: training skill-use agents
- Skills are widely used in agent harnesses, but synthesizing reliable training data and training agents to use skills remain underexplored.
- SkillGym crawls skills from the internet, keeps reproducible offline workflows, and uses a builder-reviewer pipeline to create difficulty-controlled tasks (4 task types), each with a reference solution and executable verifier.
- It yields 6.8K environments and 19K verified successful trajectories for SFT.
- Fine-tuning improves models from 2B to 122B across four skill-use benchmarks; a Qwen3.5-9B SFT model beats the untrained 397B model on two of them.
- Training raises the rate of reading the relevant skill from 28% to 96%, with gains holding across reasoning structures, minority task types, and held-out skills.
More from coding & agent
- Agent autonomously checks in a flight 24 hours ahead — and pings only on failure — armand_ruiz · 2026-10-01
- 0.8B model plus 9 LoRA adapters routes agent decisions 38x faster with +8.7 accuracy — Usual_Maximum7673 · 2026-10-01
- Context engineering's next step: knowledge system engineering — ShanRizvi · 2026-10-01
- Codex's model picker now needs a scrollbar — should OpenAI add auto-routing? — darkprinceimmortal · 2026-10-01
- OpenAI's DevDay agent launches put startups on notice, says founder in the crosshairs — NoSpecific64 · 2026-10-01
- Dev adds Apple's Genie effect to AI agent plugin with Opus, says he's basically built an agentOS — RileyRalmuto · 2026-10-01