FrogNano: a 4B model trained purely with RL on synthetic tasks hits repo-level coding
burkov · x · 2026-09-10
FrogNano shows a compact 4B model can reach strong repository-level coding performance via RL on synthetic tasks alone, without distillation from larger models. The pipeline pairs TaskPilot, an online task-synthesis system that generates executable SWE tasks from real repo snapshots calibrated to the policy's learnability frontier, with Leaf, a lightweight harness using five typed tools and a simple termination rule. Base Qwen3.5-4B is trained with group-relative DPPO over five successive rounds of task synthesis.
Related event: Microsoft's FrogNano: 4B Model Trained with Pure RL Hits 61.5% on SWE-bench(5 posts)→
More from coding & agent
- Salesforce releases EvoHarnessBench to test agents against evolving tool harnesses — Salesforce · 2026-09-10
- Tsinghua + Qwen paper: rebuilding agent workspaces beats imitating trajectories, lifts Terminal-Bench to 58.1% — rohanpaul_ai · 2026-09-10
- Users urged to make Codex/Claude document their process before models become unavailable — moonsandhues · 2026-09-10
- Whatomate: open-source single-binary WhatsApp AI chatbot platform hits 1.5k GitHub stars — tom_doerr · 2026-09-10
- Dev replaces bloated AI chat with a Kanban board that agents read and update themselves — Clean-Vermicelli-700 · 2026-09-10
- Nomos: an open-source framework for testing whether AI agent permissions have grown too broad — Excellent-Hour7253 · 2026-09-10