FrogNano trains a 4B coding agent to 61.5% SWE-bench via online task synthesis
rohanpaul_ai · x · 2026-09-21
FrogNano starts from Qwen3.5-4B and trains with RL on 1,500 synthetic software-engineering tasks, reaching 61.5% on SWE-bench Verified without frontier-model distillation. Two keys: an online curriculum that keeps generating challenging-but-learnable tasks as the agent improves, and a simpler 5-tool interface that alone lifted the base model from 8.3% to 37.2%. Paper: arxiv.org/abs/2609.07925
More from coding & agent
- 70 hands-on cybersecurity projects with full source code — tom_doerr · 2026-09-21
- Dev compares coding models building a coop game: V4.1 outshines Astra's 'pathetic' default taste — teortaxesTex · 2026-09-21
- 'Just 3 lines of code' backfires: dev argues tools should expose complexity, not hide it — willcb · 2026-09-21
- TypeSafe's Jev returns typed decisions with probabilities, not text — here's where it fits in agent loops — prakersh · 2026-09-21
- Jev Engineering gives agents a decision brain, 193x faster and 444x cheaper in tests — agihouse_org · 2026-09-21
- Dev Argues PAW Shouldn't Hide Its Complexity, Points to DSPy as the Better Playbook — willcb · 2026-09-21