20 minutes on one H200: GRPO post-training makes Qwen 3.5-2B more accurate and token-efficient

johnolafenwa · reddit · 2026-09-04

Open-source models tend to over-reason on every problem. The author released a notebook and video implementing GRPO from scratch and using it to post-train Qwen 3.5-2B for better accuracy and reasoning efficiency.

The interesting result: trained purely on simulating a Python interpreter, the model became substantially more accurate and token-efficient on math problems — a clear generalization effect. Training took 20 minutes on a single H200, and the code applies to any open-source model.

Original post →

More from Research

Research channel →