Analysis suggests Qwen 3.8 27B distilled DeepSeek-R1 via GRPO

mishig25 · x · 2026-08-15

An analysis proposes that the performance of Qwen 3.8 27B stems from distilling DeepSeek-R1-zero. The hypothesis suggests the model avoided standard SFT or massive datasets, instead using Qwen 3.8 Max's output as a reinforcement learning objective with liberal reasoning context lengths via GRPO (Group Relative Policy Optimization). This method is argued to be superior to K/L divergence distillation in certain contexts, emphasizing that intelligence is a solution, not compression.

Original post →

More from Models

Models channel →