Same GRPO recipe on three from-scratch LLMs yields wildly different results with no clean scale relationship

john_enev · reddit · 2026-08-20

The author trained three LLMs from scratch in raw PyTorch (353M/316M/672M), then applied SFT + GRPO to each with identical curriculum, reward, hyperparameters, and KL coefficient:

All nine checkpoints and a comparison playground are open on Hugging Face.

Original post →

More from Models

Models channel →