Gemma 4 Training Results Lag Behind Qwen

ivan_bezdomny · x · 2026-07-17

The author mentions that they also trained a Gemma 4 model, but its current performance isn't as good as Qwen.

They observed that one characteristic of Gemma is its tendency to generate very short responses. This has an upside—it's less prone to info-dumping like GPT-5—but the downside is that the outputs are too short for continuous learning, making training difficult. The next step will likely involve adjusting the GRPO reward.

Original post →

More from Models

Models channel →