FrogNano details: adaptive on-policy synthetic RL envs train a 4B coding model to 61.5%

sivareddyg · x · 2026-09-09

The FrogNano team shares details: Qwen3.5-4B post-trained purely with RL on synthetic TaskPilot tasks — zero distillation, zero teacher SFT, just 5 iterations × 300 tasks (stopped at 1500) — reaching 61.5% on SWE-bench Verified. The key is 'on-policy data': generating synthetic RL environments targeted at the current policy's learnability zone and adapting them continuously from rollout feedback, with surprising generalization to TB2, SwePro and PatchEval.

Related event: FrogNano: 4B Model Trained with Pure RL Hits 61.5% on SWE-bench(2 posts)→

Original post →

More from coding & agent

coding & agent channel →