Google Rolls Out Broad Gemma 4 Improvements

Multiple posts indicate that Google is pushing a sizable update to Gemma 4 based on community feedback and contributions. The reported focus areas are tool calling, chat accuracy and reliability, visual understanding, and inference speed. For users running local AI setups, tool-heavy pipelines, or agent workflows, these changes matter because they affect whether the model is more stable, faster, and better suited to longer tasks.

Key changes

Across the posts, the most repeated points are adjustments to chat templates and fixes for tool-calling issues. @Iwaku_Real and @solyarisoftware also mention reduced “laziness,” with fewer cases where the model stops too early or gives incomplete answers. On the vision side, posts say Gemma 4 now has better fine-grained image understanding and additional resources related to vision capabilities, including token budget management.

Performance and usage notes

On performance, @Iwaku_Real says Flash Attention was enabled for Gemma 4 on Hopper GPUs, while @udmrzn’s repost specifies Flash Attention 4 and says inference speed improves noticeably. @jocarrasqueira adds that the update appears to make Gemma 4 better suited to long-running agentic tasks. A repost by @danielhanchen also notes that Unsloth has advised users to re-download the updated model.

2026-07-16 ~ 2026-07-17 · 7 related posts