Gemma 4 Training Results Lag Behind Qwen
ivan_bezdomny · x · 2026-07-17
The author mentions that they also trained a Gemma 4 model, but its current performance isn't as good as Qwen.
They observed that one characteristic of Gemma is its tendency to generate very short responses. This has an upside—it's less prone to info-dumping like GPT-5—but the downside is that the outputs are too short for continuous learning, making training difficult. The next step will likely involve adjusting the GRPO reward.
More from Models
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22
- Gemini 3.5 Flash-Lite beats 3.1 Flash-Lite on long-context retrieval in MRCRv2 — Dillonu · 2026-07-22