Gemma 4 Tech Report: Open-Source Native Multimodal Models
burkov · x · 2026-07-08
Google released the Gemma 4 technical report, introducing a new generation of open-source, natively multimodal language models. The Gemma 4 series features both dense and Mixture-of-Experts (MoE) architectures, with parameter sizes ranging from 2.3B to 31B, aiming to boost computational efficiency and reasoning capabilities.
All model sizes feature improved vision and audio encoders. Notably, the 12B model adopts a unified encoder-free architecture, allowing it to process raw audio and image patches directly. Additionally, it integrates a thinking mode, enabling the model to generate a chain of thought before answering.
More from Models
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11