SkillGym: Fine-Tuning on Verified Agent Skills Lifts 35B Model Past Claude Sonnet 4.6

rohanpaul_ai · x · 2026-10-06

A Shanghai AI Lab paper, SkillGym (arXiv:2609.27717), turns human-written agent skills into sandboxed, verifiable training environments instead of inference-time instructions. It releases 2,756 environments across 12 categories and 8,364 verified successful trajectories. SFT on these runs lifts Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2 and 19.10 points on Terminal-Bench 2.1; the 35B SkillGym-Agent hits 51.47% on skill-assisted SkillsBench, exceeding reported scores of Claude Sonnet 4.6, GPT-5.4 Mini, and DeepSeek V4 Pro.

Related event: SkillGym: Fine-Tuning on Verified Skill Trajectories Boosts Agent Performance(2 posts)→

Original post →

More from coding & agent

coding & agent channel →