SkillGym: Fine-Tuning on Verified Skill Runs Lifts Terminal-Bench 2.1 Success by 19 Points
rohanpaul_ai · x · 2026-10-06
A new Shanghai AI Laboratory paper shows that training on verified runs of human-written agent skills makes a model a better agent even without the skill files.
Method
- Prompt-time skills depend on retrieval and instruction following; fine-tuning on verified skill runs reduces that dependence.
- SkillGym turns each skill file into a sandboxed task with a pass/fail code checker, then trains only on passing runs.
Results
- In Claude Code, Qwen3.5-35B-A3B jumped from 39.33% to 58.43% (+19.10 points) on Terminal-Bench 2.1.
- With no skill files loaded, it scored 26.81% on SkillsBench, beating the base model with skills (23.34%).
- Loading skills on top still helps, lifting it further to 51.47%.
Takeaway
Skill files only help when retrieved and followed; internalizing skill usage into weights is a more robust path to capable agents.
More from coding & agent
- Ethan Mollick: my AI workflow spans local Codex, cloud Dot and Claude tasks — explaining it is chaos — emollick · 2026-10-06
- Developer builds a tool that turns math equations into playable music — measure_plan · 2026-10-06
- Rork and ASC CLI to add iPhone Duo screenshot support up to 2853x2007 — rudrank · 2026-10-06
- One Prompt, $1 Budget: Agent Produces a 30s Video via Lora Pilot MCP — no3us · 2026-10-06
- Persona Policies (PPol): evolutionary framework for diverse LLM user simulators — natashajaques · 2026-10-06
- Dev builds native Mac app to message multiple AI agents simultaneously — msg · 2026-10-06