FrogNano: a 4B model trained purely with RL on synthetic tasks hits repo-level coding

burkov · x · 2026-09-10

FrogNano shows a compact 4B model can reach strong repository-level coding performance via RL on synthetic tasks alone, without distillation from larger models. The pipeline pairs TaskPilot, an online task-synthesis system that generates executable SWE tasks from real repo snapshots calibrated to the policy's learnability frontier, with Leaf, a lightweight harness using five typed tools and a simple termination rule. Base Qwen3.5-4B is trained with group-relative DPPO over five successive rounds of task synthesis.

Related event: Microsoft's FrogNano: 4B Model Trained with Pure RL Hits 61.5% on SWE-bench(5 posts)→

Original post →

More from coding & agent

coding & agent channel →