Skip SFT, run GRPO on a 0.8B Qwen base model and watch reasoning emerge

alexcovo_eth · x · 2026-09-06

A copy-paste recipe for a first RL project:

The author positions it as the ideal showcase project for anyone wanting to demo RL skills, with the models hosted on Hugging Face.

Original post →

More from coding & agent

coding & agent channel →