Microsoft's FrogNano trains a 4B coding agent almost entirely with RL, no teacher model

ossm-me · reddit · 2026-09-10

A new report called FrogNano from Microsoft Research Montréal's Froggy Team, with collaborators from Mila and UC San Diego, refines Qwen3.5-4B into a small coding agent using reinforcement learning — with no human-labeled data and no larger teacher model providing answers.

Key points:

The stated goal is a coding assistant that runs on limited hardware. It's an early report rather than a finished model, but clearly presented and carefully tested.

Related event: Microsoft's FrogNano: A 4B Coding Agent Trained with Pure RL(6 posts)→

Original post →

More from coding & agent

coding & agent channel →