FrogNano: 4B Model Trained with Pure RL Hits 61.5% on SWE-bench
FrogNano, built on Qwen3.5-4B, achieved 61.5% on SWE-bench using pure reinforcement learning on TaskPilot-synthesized tasks, with no distillation or teacher-trajectory SFT, after just 5 iterations of 300 tasks.
2026-09-09 ~ 2026-09-09 · 2 related posts
- FrogNano: 4B model hits 61.5% on SWE-bench Verified with pure RL, zero distillation — sivareddyg · 2026-09-09
- FrogNano details: adaptive on-policy synthetic RL envs train a 4B coding model to 61.5% — sivareddyg · 2026-09-09