FrogNano trains a 4B coding agent to 61.5% SWE-bench via online task synthesis

rohanpaul_ai · x · 2026-09-21

FrogNano starts from Qwen3.5-4B and trains with RL on 1,500 synthetic software-engineering tasks, reaching 61.5% on SWE-bench Verified without frontier-model distillation. Two keys: an online curriculum that keeps generating challenging-but-learnable tasks as the agent improves, and a simpler 5-tool interface that alone lifted the base model from 8.3% to 37.2%. Paper: arxiv.org/abs/2609.07925

Original post →

More from coding & agent

coding & agent channel →