SkillGym: Fine-Tuning on Verified Agent Skills Lifts 35B Model Past Claude Sonnet 4.6
rohanpaul_ai · x · 2026-10-06
A Shanghai AI Lab paper, SkillGym (arXiv:2609.27717), turns human-written agent skills into sandboxed, verifiable training environments instead of inference-time instructions. It releases 2,756 environments across 12 categories and 8,364 verified successful trajectories. SFT on these runs lifts Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2 and 19.10 points on Terminal-Bench 2.1; the 35B SkillGym-Agent hits 51.47% on skill-assisted SkillsBench, exceeding reported scores of Claude Sonnet 4.6, GPT-5.4 Mini, and DeepSeek V4 Pro.
More from coding & agent
- Used OpenAI Dots as a Free Agent Swarm to Break a 47-Year-Old Math Record — jaxchang · 2026-10-06
- Build an Internal Company Chatbot by Wiring Grok Bot to Slack and Your SOPs — alexcovo_eth · 2026-10-06
- Dev shares semantic router results: guard model catches 87% of attacks — colinmcnamara · 2026-10-06
- Graphical launches: a design tool for visual languages that pairs with coding agents — mikelikesdesign · 2026-10-06
- Context Compression Hurts Long-Horizon Agents at a Few Key Points, PAIR Pinpoints Why — dair_ai · 2026-10-06
- Dev ships free Cloudflare Worker that lets agents email you only when it matters — CyrisXD · 2026-10-06