Gemma 4 31B can use over 13× more KV-cache memory than DeepSeek V4 Flash
teortaxesTex · x · 2026-08-04
- The post compares KV-cache memory costs for Gemma 4 31B and DeepSeek V4 Flash using a cache-size calculator screenshot.
- For Gemma 4 31B, the screenshot shows 38.54 GiB total cache at 1M tokens, with 40.41 KiB per token and 320 GB cache at batch=32 / 256K sequence length mentioned in the text.
- The author argues Gemma uses 2.4× more active parameters and 13.3× more bits of memory per token than DeepSeek at equal KV precision.
- The DeepSeek screenshot shows a much smaller 2.89 GiB total cache under the same 1M-token setup, illustrating how architecture choices affect inference memory economics.
More from Infra
- China’s AI chip shipments seen rising from 3.8M to 12.9M by 2028 — teortaxesTex · 2026-08-04
- Long-context voice agents ditch turn detection with async compaction handoff — juberti · 2026-08-04
- MiniMax H3 lands day-one local serving on 2× RTX 5090s or 1× RTX Pro 6000 — yvbbrjdr · 2026-08-04
- MiniMax H3 video generation on ComfyUI needs 14–15 GB VRAM and over an hour at 1280×704 — ayakitodev · 2026-08-04
- Atomic releases 14 DeepSeek-V4-Flash GGUF quants, recommends AD-IQ2_M for 128GB rigs — testingcatalog · 2026-08-04
- Quoted post sees a two-year melt-up coming for neoclouds — abhiadesai · 2026-08-04