SkillGym fine-tuning lifts Qwen3.5 35B past Claude Sonnet 4.6 on agentic coding benchmarks
dair_ai · x · 2026-09-28
dair-ai highlights a paper exploring continual training of specialized models on skills via SkillGym:
- Human-written skills are turned into 2,756 training environments with code-based checkers, then the model is fine-tuned on 8,364 successful trajectories.
- After fine-tuning, Qwen3.5-35B-A3B running in Claude Code gains 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, reaching 51.47%.
- That exceeds reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini.
- Notably, without skills loaded, the trained model still beats the base model running with skills in context — the skills are internalized.
More from coding & agent
- Creator hooks Claude to MCP and lets it autonomously edit motion graphics — aziz4ai · 2026-09-28
- Writer lets Codex decide his book's margins, printing test to follow — santiviquez · 2026-09-28
- kindling runs every MCP server in its own Firecracker microVM, frozen at 0 RAM until called — Ascci52 · 2026-09-28
- PurzBeats demos Sentinel, an agent framework that builds its own nodes — PurzBeats · 2026-09-28
- Jev seen as fit for high-level robotics orchestration: fast schema beats LLM lag — ai · 2026-09-28
- Apple's free on-device fm paired with decision model Jev beats big-model routing in tests — jasonkneen · 2026-09-28