FrogNano details: adaptive on-policy synthetic RL envs train a 4B coding model to 61.5%
sivareddyg · x · 2026-09-09
The FrogNano team shares details: Qwen3.5-4B post-trained purely with RL on synthetic TaskPilot tasks — zero distillation, zero teacher SFT, just 5 iterations × 300 tasks (stopped at 1500) — reaching 61.5% on SWE-bench Verified. The key is 'on-policy data': generating synthetic RL environments targeted at the current policy's learnability zone and adapting them continuously from rollout feedback, with surprising generalization to TB2, SwePro and PatchEval.
Related event: FrogNano: 4B Model Trained with Pure RL Hits 61.5% on SWE-bench(2 posts)→
More from coding & agent
- ChatGPT iOS update brings async question tool support to Remote — Dimillian · 2026-09-09
- Agent Client Protocol pitches local AI agents as an escape from data handover — RileyRalmuto · 2026-09-09
- A Tool-Agnostic 3-Step Framework for Evaluating AI Agents in 2026 — Al_Grigor · 2026-09-09
- Agent shipped 10 features in a day, all 374 tests green — the page didn't render — Sinjared · 2026-09-09
- He built a playable multiplayer racing game in 2 hours on his phone — and shares the two prompts that did it — blakesamic · 2026-09-09
- 15-year video editor saves 10 hours a day editing with Claude + DaVinci MCP — _AustinCalvert_ · 2026-09-09